ElevenLabs Logo

ElevenLabs

Research Engineer - Inference

Posted 5 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in United States
Entry level
Remote
Hiring Remotely in United States
Entry level
Deploy and optimize frontier AI models for real-time production use at scale. Responsibilities include improving inference latency, throughput, and cost through quantization, distillation, KV-cache optimization, batching, and custom kernels; building high-performance streaming serving systems; and creating tooling that enables safe, efficient model deployment.
The summary above was generated by AI
About ElevenLabs

ElevenLabs is an AI research and product company transforming how we interact with technology.

We launched in January 2023 with the first human-like AI voice model. Today, we serve millions of users and thousands of businesses - from fast-growing startups to large enterprises like Deutsche Telekom and Meta. Our investors are some of the world's most prominent, including Andreessen Horowitz, ICONIQ Growth and Sequoia. We've raised $781M in funding and our last valuation was $11B - multiples of 11, always.
We have expanded from voice into three main platforms:

  • ElevenAgents enables businesses to deliver seamless and intelligent customer experiences, with the integrations, testing, monitoring, and reliability necessary to deploy voice and chat agents at scale.

  • ElevenCreative empowers creators and marketers to generate and edit speech, music, image, and video across 70+ languages.

  • ElevenAPI gives developers access to our leading AI audio foundational models.

Everything we do is the result of the creativity and commitment of our team - builders doing the best work of their lives. We are researchers, engineers, and operators. IOI medalists and ex-founders. If you want to work hard and create lasting positive impact, we want to hear from you.

How we work
  • High-velocity: Rapid experimentation, lean autonomous teams, and minimal bureaucracy.

  • Impact not job titles: We don’t have job titles. Instead, it’s about the impact you have. No task is above or beneath you.

  • AI first: We use AI to move faster with higher-quality results. We do this across the whole company—from engineering to growth to operations.

  • Excellence everywhere: Everything we do should match the quality of our AI models.

  • Global team: We prioritize your talent, not your location.

What we offer
  • Innovative culture: You’ll be part of a generational opportunity to define the trajectory of AI, surrounded by a team pushing the boundaries of what’s possible.

  • Growth paths: Joining ElevenLabs means joining a dynamic team with countless opportunities to drive impact - beyond your immediate role and responsibilities.

  • Learning & development: ElevenLabs proactively supports professional development through an annual discretionary stipend.

  • Social travel: We also provide an annual discretionary stipend to meet up with colleagues each year, however you choose.

  • Annual company offsite: Each year, we bring the entire team together in a new location - past offsites have included Croatia and Italy.

  • Co-working: If you’re not located near one of our main hubs, we offer a monthly co-working stipend.

About the role

We are looking for a Research Engineer to join the research team at ElevenLabs, focused on deploying and optimizing our frontier AI models in production. The quality of our models only matters if they can be served fast, reliably, and at scale. You will own the systems that turn research breakthroughs into real-time products used by millions. You will thrive in this role if you enjoy:

  • Deploying state-of-the-art models to production and owning the path from research checkpoint to serving infrastructure.

  • Optimizing inference performance across the stack, including latency, throughput, and cost, using techniques such as quantization, distillation, KV-cache optimization, batching strategies, and custom kernels.

  • Building and tuning high-performance serving systems for real-time, streaming workloads where every millisecond matters.

  • Creating tooling and infrastructure that lets researchers ship new models to production quickly, safely, and with confidence in their performance characteristics.

Requirements

We do not require any formal certifications or degrees. Instead, we are seeking enthusiastic engineers who can showcase solving impressively hard problems with artifacts such as past projects, designs, or GitHub contributions. Ideally, you bring:

  • Experience deploying and serving ML models in production, ideally for latency-sensitive or real-time applications.

  • Strong engineering skills in GPU programming and inference optimization (e.g., CUDA, Triton, TensorRT, or serving frameworks such as vLLM or SGLang).

  • The capacity to autonomously profile, diagnose, and eliminate bottlenecks across the serving stack, from model architecture to kernels to orchestration, and to build the tooling to measure it.

Location

This role is remote and can be executed globally. If you prefer, you can work from our offices in London, New York, San Francisco, and Warsaw.

#LI-Remote

We are an equal opportunity employer and do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, veteran status, disability or other legally protected statuses.

Similar Jobs

2 Hours Ago
Easy Apply
Remote
Easy Apply
328K-492K Annually
Expert/Leader
328K-492K Annually
Expert/Leader
Cloud • Security • Software • Cybersecurity • Automation
Provides cross-team technical leadership for large backend initiatives, modular architecture, monolith modernization, API and platform standards, production operations, and AI-assisted engineering adoption. Designs scalable services and shared patterns, improves security, reliability, observability, and operability, participates in incident response, partners with product and engineering leadership, and mentors engineers.
Top Skills: Ai-Assisted Engineering WorkflowsAPIsDistributed SystemsEvent-Driven ArchitectureGoKubernetesPostgresRuby On RailsRust
2 Hours Ago
Easy Apply
Remote
Easy Apply
248K-372K Annually
Mid level
248K-372K Annually
Mid level
Cloud • Security • Software • Cybersecurity • Automation
Build, test, ship, and operate production backend features with minimal guidance. Design and maintain APIs, background jobs, integrations, data models, migrations, indexes, and service boundaries. Improve security, performance, reliability, observability, and maintainability; investigate production issues; participate in code reviews and operational support. Collaborate asynchronously with Product Management, Frontend, Infrastructure, and Security, while using and validating AI coding tools.
Top Skills: Ai Coding ToolsAPIsCi/CdDockerGoGrpcHelmKubernetesPostgresRuby on RailsRubyRust
6 Hours Ago
Easy Apply
Remote
Easy Apply
Senior level
Senior level
Big Data • Fintech • Mobile • Payments • Financial Services
Own quarterly engineering goals for the Recoveries team, delivering scalable backend systems that maximize charged-off loan recovery and reconcile partner discrepancies. Lead engineers through ambiguous problems, collaborate with product, design, analytics, and risk stakeholders, and ensure operational reliability through metrics, monitoring, on-call support, and failure handling. Establish quality standards, guide technical design and implementation, contribute to large codebases, and mentor engineers through feedback and technical leadership.
Top Skills: AWSKotlinKubernetesMySQLPythonReactVue

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account