AHEAD Logo

AHEAD

Data Engineer

Posted Yesterday
Remote
Hiring Remotely in United States
150K-180K Annually
Mid level
Remote
Hiring Remotely in United States
150K-180K Annually
Mid level
Build and operate cloud data capabilities, including ingestion pipelines, transformations, data models, curated data products, and data-quality controls. Use Snowflake, dbt, SQL, and Python to deliver governed data for analytics, applications, automation, and AI workflows. Implement testing, CI/CD, documentation, lineage, access controls, monitoring, incident resolution, and performance optimization. Collaborate with engineering, analytics, governance, security, and business teams in an Agile, product-oriented environment with active AI-assisted development.
The summary above was generated by AI

The Data Engineer, Data Platform will build and operate the data capabilities that help AHEAD teams access trusted, usable, and well-managed information. This role will develop ingestion pipelines, transformations, data models, and curated data products in the modern cloud data platform, with an emphasis on Snowflake and dbt. 

The role will support data coming from enterprise applications and services, including Salesforce, Hatch, NetSuite, Signal, and approved APIs. The Data Engineer will help make data available for analytics, applications, automation, and AI-enabled workflows through consistent engineering patterns,documented definitions, appropriate access controls, and dependable operational practices. Active use of AI throughout the software development lifecycle is a core expectation of this role, including AI-assisted code generation, automated testing, documentation, troubleshooting, and review with appropriate human validation. 

Working under the Director, Data Platform and alongside the Data Governance Lead, this role will contribute to a product-oriented engineering team. The role will partner with data consumers and other engineering teams to understand requirements, deliver useful platform capabilities, and improve the speed and consistency of data delivery. 

Duties/Responsibilities

  • Build, maintain, and improve batch and low-latency data ingestion pipelines from enterprise systems, APIs, and other approved sources. 
  • Follow the AI SDLC by actively using approved AI coding tools and agents to generate, refactor, explain, and review code; validate generated output through engineering judgment, testing, and peer review. 
  • Use AI to generate and improve unit, integration, data-quality, and regression tests, then verify that automated tests accurately validate the intended behavior. 
  • Use AI-assisted workflows to create and maintain technical documentation, data-product documentation, runbooks, lineage notes, and change summaries as part of delivery. 
  • Build toward coordinated multi-agent delivery patterns that can divide and accelerate discovery, implementation, testing, documentation, and operational support while preserving human accountability. 
  • Develop SQL and Python solutions that collect, validate, transform, and publish data for downstream consumption. 
  • Use Snowflake and dbt to implement reliable transformations, reusable models, curated datasets, and data products across raw, common, and curated layers. 
  • Translate business and technical requirements into source mappings, data models, acceptance criteria, and maintainable engineering solutions. 
  • Partner with analytics, application, AI, Integration Platform, and business teams to make data available through governed and documented access patterns. 
  • Apply data quality checks for completeness, freshness, uniqueness, consistency, referential integrity, and other relevant quality dimensions. 
  • Add metadata, documentation, lineage, ownership, and usage guidance to data products so consumers can find and understand the data they use. 
  • Implement secure access patterns in partnership with Data Governance and Security teams, including role-based access, classification tags, masking, and row- or column-level controls when appropriate. 
  • Build automated tests and deployment processes that support consistent delivery through development, quality assurance, and production environments. 
  • Monitor pipeline health, data freshness, processing performance, and failures; troubleshoot issues and participate in incident resolution. 
  • Optimize Snowflake workloads, queries, transformations, and storage patterns for performance, reliability, and cost discipline. 
  • Support the curation and publication of cross-system data needed for shared business context, entity-aware access, reporting, automation, and AI use cases. 
  • Work with the Integration Platform and semantic-layer capabilities, including Horizon, to support consistent business meaning and reusable data access. 
  • Participate in backlog refinement, estimation, code review, technical documentation, and iterative delivery within an Agile engineering team. 
  • Identify opportunities to simplify delivery, reduce duplicate work, improve platform standards, and strengthen the reliability of data engineering practices. 

Education and Experience

    • Bachelor’s degree in computer science, information systems, engineering, mathematics, or a related field, or equivalent experience. 

    • 3 or more years of experience in data engineering, software engineering, analytics engineering, or a related technical role. 

    • Professional experience writing production-quality SQL and Python. 

    • Experience building or supporting data pipelines, transformations, and data models in a cloud data environment. 

    • Experience with Snowflake, dbt, or comparable cloud data warehouse and transformation technologies. 

    • Understanding of data modeling, ELT/ETL patterns, pipeline orchestration, APIs, and source-system integration. 

    • Experience with software engineering practices including source control, code review, automated testing, and CI/CD. 

    • Demonstrated active use of AI-assisted software development tools for code generation, test creation, documentation, debugging, or review. 

    • Ability to follow an AI SDLC and identify practical opportunities for multiple cooperating agents to improve delivery speed, consistency, and coverage. 

    • Understanding of data quality, metadata, lineage, access control, privacy, and secure handling of enterprise data. 

    • Ability to investigate data issues, communicate findings clearly, and work through ambiguity with teammates and stakeholders. 

    • Ability to collaborate effectively with engineers, analysts, product owners, governance partners, security teams, and business stakeholders. 

Preferred

    • Experience with Azure services, serverless functions, cloud storage, or other cloud-native data engineering capabilities. 

    • Experience with REST or GraphQL APIs and data ingestion from enterprise applications such as Salesforce, Hatch, NetSuite, or similar systems. 

    • Familiarity with orchestration, event-driven processing, observability, data catalogs, lineage tooling, or data quality platforms. 

    • Experience supporting semantic models, MCP-based access, or other governed interfaces for analytics, applications, automation, or AI workflows. 

    • Experience working with master data, reference data, entity resolution, or shared business definitions across multiple systems. 

    • Experience operating data products with documented ownership, access expectations, quality measures, and support procedures. 

    • Experience using AI agents or agentic workflows to support software delivery, data engineering, testing, documentation, or platform operations. 

    • Curiosity about emerging data platform technologies and a practical approach to adopting them. 

     

Physical Requirements

     
  • Ability to safely and successfully perform the essential job functions consistent with the ADA, FMLA, and other federal, state, and local standards, including meeting qualitative and/or quantitative productivity standards. 
    • Ability to maintain regular, punctual attendance consistent with the ADA, FMLA, and other federal, state, and local standards. 

    • Primarily office and computer-based work with standard engineering and collaboration expectations for an enterprise technology role. 

Similar Jobs

7 Days Ago
Easy Apply
Remote or Hybrid
United States
Easy Apply
Senior level
Senior level
Fintech • News + Entertainment • Software • Database • Financial Services
Lead the architecture and development of scalable AWS-based data ingestion, transformation, and orchestration pipelines. Build reliable data infrastructure using Python, SQL, Airflow, Lambda, ECS, SQS, and Terraform. Establish data modeling, quality, lineage, monitoring, testing, and observability practices while partnering with analysts, scientists, and backend engineers. Mentor senior engineers, guide technical decisions, and provide hands-on leadership for complex data platform initiatives.
Top Skills: Amazon EcsAmazon KinesisAmazon RedshiftAmazon S3Amazon SqsApache AirflowApache FlinkAWSAws GlueAws LambdaBeautifulsoupCi/CdDatabricksDockerGreat ExpectationsKafkaMonte CarloMwaaPythonScrapySnowflakeSQLTerraform
12 Days Ago
Remote or Hybrid
113K-193K Annually
Senior level
113K-193K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Design and operate scalable batch and streaming data platforms supporting machine learning and generative AI. Build pipelines for structured, unstructured, OCR, document, and image data; develop RAG, semantic search, and LLM-powered solutions; and establish data quality, observability, governance, orchestration, and deployment practices. Partner with stakeholders, lead platform scalability and cost optimization, mentor engineers, and translate ambiguous needs into production-ready technical roadmaps while securely handling sensitive data.
Top Skills: AirflowAmazon KinesisSparkAWSAzureAzure Event HubsChart.JsDatabricksDeequDelta LakeDockerGithub ActionsGCPGreat ExpectationsJavaKafkaKubernetesLlmsMlopsPlotlyPysparkPythonRagScalaSeabornSnowflakeSQLTerraform
13 Days Ago
Remote
United States
Mid level
Mid level
Artificial Intelligence • Information Technology • Professional Services • Software • Analytics • Generative AI • Big Data Analytics
Configure and maintain Adobe Experience Platform and Real-Time CDP data pipelines, XDM schemas, datasets, identity resolution, Profile enablement, segmentation readiness, and destination activation. Validate data quality, troubleshoot ingestion and identity issues, manage sandboxes, and implement privacy, consent, governance, and access controls. Partner with architects, data engineers, consultants, analysts, data scientists, and marketing teams to support reporting, personalization, and machine learning readiness.
Top Skills: Adobe AnalyticsAdobe Experience PlatformAdobe I/O RuntimeAdobe Journey OptimizerAdobe Real-Time CdpAdobe Source ConnectorsAdobe TargetAPIsAWSAzureBatch IngestionBigQueryCcpaData PrepEltETLGCPGdprJavaScriptPythonQuery ServiceRedshiftSalesforce CdpSegmentSnowflakeSQLStreaming IngestionWeb SdkXdm

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account