People Data Labs Logo

People Data Labs

Data Engineer

Reposted 11 Days Ago
Remote
Hiring Remotely in USA
160K-180K
Mid level
Remote
Hiring Remotely in USA
160K-180K
Mid level
The Data Engineer will build data infrastructure, develop CI/CD pipelines, and create solutions for data engineering challenges, requiring strong technical skills and collaboration.
The summary above was generated by AI

Note for all engineering roles: with the rise of fake applicants and AI-enabled candidate fraud, we have built in additional measures throughout the process to identify such candidates and remove them.

About Us

People Data Labs (PDL) is the provider of people and company data. We do the heavy lifting of data collection and standardization so our customers can focus on building and scaling innovative, compliant data solutions. Our sole focus is on building the best data available by integrating thousands of compliantly sourced datasets into a single, developer-friendly source of truth. Leading companies across the world use PDL’s workforce data to enrich recruiting platforms, power AI models, create custom audiences, and more.

We are looking for individuals who can balance extreme ownership with a “one-team, one-dream” mindset. Our customers are trying to solve complex problems, and we only help them achieve their goals as a team. Our Data Engineering Team is the secret sauce behind all that we do and we are looking for the best of the best.

If you are looking to be part of a team discovering the next frontier of data-as-a-service (DaaS) with a high level of autonomy and opportunity for direct contributions, this might be the role for you. We like our engineers to be thoughtful, quirky, and willing to fearlessly try new things. Failure is embraced at PDL as long as we continue to learn and grow from it.

What You Get to Do

  • Build infrastructure for ingestion, transformation, and loading an exponentially increasing volume of data from a variety of sources using Spark, SQL, AWS, and Databricks
  • Building an organic entity resolution framework capable of correctly merging hundreds of billions of individual entities into a number of clean, consumable datasets.
  • Developing CI/CD pipelines and anomaly detection systems capable of continuously improving the quality of data we're pushing into production.
  • Dreaming up solutions to largely undefined data engineering and data science problems.

The Technical Chops You’ll Need

  • 4-6+ years of industry experience with clear examples of strategic technical problem-solving and implementation
  • Strong software development fundamentals.
  • Experience with Python 
  • Expertise with Apache Spark (Java, Scala, and/or Python-based)
  • Experience with SQL
  • Experience building scalable data processing systems (e.g., cleaning, transformation)  from the ground up.
  • Experience using developer-oriented data pipeline and workflow orchestration (e.g., Airflow (preferred), dbt, dagster or similar)
  • Knowledge of modern data design and storage patterns (e.g., incremental updating, partitioning and segmentation, rebuilds and backfills)
  • Experience working in Databricks (including delta live tables, data lakehouse patterns, etc.)
  • Experience with cloud computing services (AWS (preferred), GCP, Azure or similar)
  • Experience with data warehousing (e.g., Databricks, Snowflake, Redshift, BigQuery, or similar)
  • Understanding of modern data storage formats and tools (e.g., parquet, ORC, Avro, Delta Lake)

People Thrive Here Who Can

  • Balance high ownership and autonomy with a strong ability to collaborate
  • Work effectively remotely (able to be proactive about managing blockers, proactive on reaching out and asking questions, and participating in team activities)
  • Demonstrate strong written communication skills on Slack/Chat and in documents
  • Exhibt experience in writing data design docs (pipeline design, dataflow, schema design)
  • Scope and breakdown projects, communicate and collaborate progress and blockers effectively with your manager, team, and stakeholders

Some Nice To Haves

  • Degree in a quantitative discipline such as computer science, mathematics, statistics, or engineering
  • Experience working with entity data (entity resolution / record linkage)
  • Experience working with data acquisition / data integration
  • Expertise with Python and the Python data stack (e.g., numpy, pandas)
  • Experience with streaming platforms (e.g., Kafka)
  • Experience evaluating data quality and maintaining consistently high data standards across new feature releases (e.g., consistency, accuracy, validity, completeness)

Our Benefits

  • Stock
  • Competitive Salaries
  • Unlimited paid time off
  • Medical, dental, & vision insurance 
  • Health, fitness, and office stipends
  • The permanent ability to work wherever and however you want

Comp: $160-180K

People Data Labs does not discriminate on the basis of race, sex, color, religion, age, national origin, marital status, disability, veteran status, genetic information, sexual orientation, gender identity or any other reason prohibited by law in provision of employment opportunities and benefits.

Qualified Applicants with arrest or conviction records will be considered for Employment in accordance with the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act.

Personal Privacy Policy for California Residents

https://www.peopledatalabs.com/pdf/privacy-policy-and-notice.pdf

Top Skills

Airflow
Spark
Avro
AWS
Azure
BigQuery
Databricks
Databricks
Dbt
Delta Lake
GCP
Orc
Parquet
Python
Redshift
Snowflake
Spark
SQL

Similar Jobs

3 Days Ago
In-Office or Remote
2 Locations
Senior level
Senior level
Artificial Intelligence • Enterprise Web • Machine Learning • Natural Language Processing • Software • Conversational AI • Automation
As a Senior Data Engineer, you'll enhance data architecture and streamline data ingestion processes to improve product offerings and insights. Your responsibilities include data analysis, debugging, production support, and consulting on best practices for data management and architecture.
Top Skills: AWSBigQueryDynamoDBElasticsearchGCPGoIcebergJavaScriptKinesisMongoDBPythonS3StarrocksTypescript
9 Days Ago
Remote or Hybrid
Orlando, FL, USA
Senior level
Senior level
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
As a Data Engineer II, you will design, build, and maintain data pipelines, troubleshoot production issues, and work in Agile teams to support Fandango's data needs.
Top Skills: Amazon MwaaApache AirflowAthenaAWSDynamoDBEmrGlueGlueInformaticaJavaLambdaPysparkPythonRedshiftS3ScalaSQLTalendTerraform
10 Days Ago
Remote or Hybrid
New York, NY, USA
150K-180K Annually
Senior level
150K-180K Annually
Senior level
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Lead design and development of SAP data solutions, ensuring data quality, integrity, and compliance while collaborating within agile teams.
Top Skills: AWSAzureDatabricksGCPPythonRest ApisSap Analytics CloudSap Business Data CloudSap BwSap Bw/4HanaSap DatasphereSap Hana

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account