Wynd Labs Logo

Wynd Labs

Machine Learning Engineer

Reposted One Month Ago
Remote
Hiring Remotely in USA
Mid level
Remote
Hiring Remotely in USA
Mid level
Develop, fine-tune, and deploy LLMs for NLP tasks; design and maintain large-scale data pipelines; analyze time-series data; implement data-driven models and OCR systems; ensure data quality; collaborate cross-functionally; and research and apply best practices to improve internal data processing tools and infrastructure.
The summary above was generated by AI

Who We Are:

We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models.

We're the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs.

We’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI.

The Role:

We are looking for a Machine Learning Engineer with strong skills and significant experience developing machine learning models. You will join a small, innovative team and lead efforts to advance our capabilities, drive model development, and support our vision for a future where Grass is transformative in the internet's evolution.

Please note: This role requires a work schedule that sufficiently overlaps with EST business hours to collaborate effectively with the team.

Who You Are:

  • Bachelor’s, Master’s, or Doctoral degree in Data Science, Computer Science, Statistics, or a related field.

  • A minimum of 3 years of work or research experience dealing with large datasets.

  • Experience working with large-scale text datasets, NLP pipelines, or data preparation for LLM training is highly preferred.

  • Strong coding skills in Python or other object-oriented programming languages.

  • Experience with text deduplication, dataset filtering, corpus curation, or data distillation is a strong plus.

  • Graduate-level knowledge of statistics, including but not limited to hypothesis testing, regression analysis, and probability.

  • Excellent work ethic and the ability to thrive in a fast-paced startup environment.

  • Strong problem-solving skills and attention to detail.

  • Good communication skills, with the ability to articulate complex data concepts to non-technical stakeholders.

  • Experience working in a high-output team.

What You'll Be Doing:

  • Developing data processing pipelines and machine learning solutions for large-scale NLP and LLM applications, including improving the quality, filtering, and preparation of training datasets.

  • Designing and implementing pipelines for processing and analyzing large datasets.

  • Analyzing and interpreting complex time series data to provide actionable insights and solutions.

  • Designing, implementing, and maintaining data-driven models and algorithms.

  • Developing techniques for dataset curation to improve the quality and efficiency of AI training data.

  • Building scalable pipelines for filtering, deduplicating, and improving large-scale text datasets used for LLM training.

  • Collaborating with cross-functional teams to understand data needs and deliver timely solutions.

  • Ensuring data quality and integrity throughout all processes.

  • Utilizing Optical Character Recognition (OCR) technology to convert different types of documents into editable and searchable data.

  • Continuously researching and implementing best practices in data science and machine learning.

  • Contributing to the development and improvement of internal data processing tools and infrastructure.

Why Work With Us:

  • Opportunity. We are at the forefront of developing a web-scale crawler and knowledge graph that improves access to public web data and extends the value of AI to the people.

  • Culture. We're a lean team with a high bar. We come to work not to be comfortable, but to find out what we're capable of and to do work that matters. We're not calling for people who keep things moving. We're calling for people who make everyone around them better.
    We prioritize low ego and high output. This is a fully remote team.

  • Compensation. You’ll receive a competitive salary, benefits and equity package.

Similar Jobs

Yesterday
Easy Apply
Remote or Hybrid
Easy Apply
158K-225K Annually
Senior level
158K-225K Annually
Senior level
Cloud • Information Technology • Security • Software • Cybersecurity
Build and operate production LLM-powered agents and machine learning systems for customer risk identification and threat detection. Translate security heuristics into agent logic, collaborate with threat researchers, develop data requirements and pipelines, evaluate precision and recall, and deploy solutions through CI/CD. The role also requires production reliability ownership, observability, incident debugging, cloud infrastructure expertise, and senior staff-level technical leadership.
Top Skills: AthenaAWSCi/CdDockerElasticsearchLangchainLanggraphLlm AgentsNumpyOpensearchPandasPolarsPrestoPythonSQL
2 Days Ago
In-Office or Remote
120K-215K Annually
Senior level
120K-215K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Build and deploy enterprise agentic AI solutions, including agents, MCP tools, plugins, orchestration architectures, evaluation harnesses, and secure APIs. Integrate agents with search and collaboration platforms, instrument production behavior, manage cost and latency, and test against prompt injection and tool vulnerabilities. The role requires backend software engineering, production agent development, MCP experience, LLM evaluation, and AI coding assistant expertise, with preferred experience in managed runtimes, retrieval systems, and enterprise security.
Top Skills: Ai/MlAksAPIsAws StrandsAzure Ai Foundry Agent ServiceAzure Ai SearchBackend Software EngineeringBedrock AgentcoreBot FrameworksBraintrustCi/CdClaude Agent SdkClaude CodeCodexCopilotDeepevalEksEmbeddingsGkeGleanGoogle AdkLangfuseLanggraphLangsmithMicrosoft TeamsModel Context Protocol (Mcp)Openai Agents SdkOpensearchPgvectorPostgresPromptfooPydantic AiPythonRagRagasRanxRbacSecrets ManagementService IntegrationSlackSsoVector StoresVertex Ai Agent Engine
2 Days Ago
In-Office or Remote
146K-250K Annually
Senior level
146K-250K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Lead development and deployment of an agentic AI/ML pipeline for identifying, prioritizing, and pricing high-cost healthcare claims. Responsibilities include model development, confidence scoring, HITL safety design, ML lifecycle management, monitoring, workflow automation, responsible AI, and enterprise-scale analytics using Databricks and Snowflake. The role also evaluates emerging AI technologies and supports regulatory-aware automation and healthcare data transformation.
Top Skills: Agentic Ai OrchestrationAi/MlContinuous IntegrationDatabricksDatabricks Medallion ArchitectureEnsemble ModelsHuman-In-The-Loop (Hitl)Machine Learning Lifecycle ManagementMlopsModel DeploymentModel MonitoringOpentelemetryPythonSnowflake

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account