Wynd Labs Logo

Wynd Labs

Web Scraping Specialist

Reposted 3 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Lead development and optimization of web scraping pipelines to extract, clean, and store large-scale web data. Handle dynamic content, pagination, distributed scraping, database design with NoSQL, deploy jobs to cloud, and monitor systems for reliability and data quality. Support ML-based data cleaning and categorization.
The summary above was generated by AI

Who We Are:

We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models.

We're the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs.

We’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI.

The Role.

We are seeking a Web Scraping Specialist who is proficient and brings significant experience in data extraction and web scraping techniques. You will join a small, specialized team and lead efforts to gather and analyze data, optimize scraping processes, and support our vision for a future where Grass plays a crucial role in transforming internet data accessibility.

Who You Are.

  • Demonstrated ability to extract data from complex websites with minimal supervision, with a portfolio or examples of past projects.

  • Proficiency in languages such as Python or JavaScript, with strong skills in libraries and frameworks like BeautifulSoup, Scrapy, or Selenium.

  • Knowledge of asynchronous programming, multithreading, and distributed scraping.

  • In-depth knowledge of HTML, CSS, JavaScript, and the Document Object Model (DOM).

  • Experience with NoSQL databases (MongoDB, Cassandra), capable of designing efficient storage solutions and managing data integrity.

  • Ability to apply machine learning algorithms for data cleaning, categorization, or predictive analysis adds significant value.

  • Experience with cloud services (AWS, Google Cloud, Azure) for deploying and managing scraping jobs at scale.

  • Active participation in open-source projects related to web scraping, data processing, or similar fields.

What You'll Be Doing.

  • Write, test, and refine code that extracts data from various online sources, ensuring reliability and efficiency.

  • Perform data retrieval tasks, handling complexities such as pagination and dynamic content loaded with AJAX.

  • Clean and format extracted data, ensuring it meets quality standards for further analysis or processing.

  • Database management: Store and manage the scraped data in appropriate databases, optimizing for access speed and data integrity.

  • Regularly monitor the scraping processes, identify and resolve any issues to maintain continuous data flow.

Why Work With Us:

  • Opportunity. We are at the forefront of developing a web-scale crawler and knowledge graph that improves access to public web data and extends the value of AI to the people.

  • Culture. We're a lean team with a high bar. We come to work not to be comfortable, but to find out what we're capable of and to do work that matters. We're not calling for people who keep things moving. We're calling for people who make everyone around them better.
    We prioritize low ego and high output. This is a fully remote team.

  • Compensation. You’ll receive a competitive salary, benefits and equity package.

Similar Jobs

17 Days Ago
In-Office or Remote
2 Locations
75K-100K Annually
Mid level
75K-100K Annually
Mid level
Artificial Intelligence • Blockchain • Information Technology • Consulting
Build and maintain high-performance web scraping pipelines: develop extraction code, handle dynamic content and pagination, ensure data quality and storage, monitor and scale distributed scraping infrastructure, and optimize processes for large-scale AI data ingestion.
Top Skills: Asynchronous ProgrammingAWSAzureBeautifulsoupCassandraCSSDistributed Scraping ArchitecturesDomGCPHTMLJavaScriptMachine LearningMongoDBMultithreadingPythonScrapySelenium
22 Minutes Ago
Remote
United States
120K-220K Annually
Entry level
120K-220K Annually
Entry level
Fintech • Information Technology • Software
Build, train, and productionize machine learning models for fraud detection and identity verification. Perform data acquisition, feature engineering, experimentation, monitoring, and analyses to inform product, sales, and risk decisions. Collaborate with engineering, data acquisitions, and risk operations to maintain data quality and ship production-ready code. Research new fraud types and contribute to new product development across the ML lifecycle.
Top Skills: AWSEc2PostgresPythonRdsRedshiftS3
22 Minutes Ago
Remote
United States
120K-220K Annually
Entry level
120K-220K Annually
Entry level
Fintech • Information Technology • Software
Build and deploy production ML models for fraud and identity risk across the full ML lifecycle. Research new fraud types, engineer features, write production-ready code, collaborate with cross-functional teams, and inform product, data acquisition, and business decisions.
Top Skills: Aws Ec2Aws RdsAws RedshiftAws S3PostgresPython

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account