OpenRouter Logo

OpenRouter

Site Reliability Engineer, Provider Operations

Posted Yesterday
Remote
Hiring Remotely in US
Mid level
Remote
Hiring Remotely in US
Mid level
Own the operational health, observability, reliability, and failover of AI model providers and endpoints. Build SLO monitoring, canaries, quality regression detection, automated provider-operations tooling, and load-testing systems. Lead incident response, communicate with external providers, produce reliability scorecards, and drive postmortems and corrective actions. Support capacity planning and launch readiness for high-volume LLM inference traffic.
The summary above was generated by AI
About OpenRouter

OpenRouter is the leading AI routing and infrastructure layer that enterprises use to access, manage, and optimize the best large language models across providers—without lock-in, capacity constraints, or unnecessary cost. We power the most advanced AI teams in the world by giving them the flexibility to move fast, scale confidently, and stay future-proof as models evolve.

As enterprise adoption of AI accelerates, OpenRouter sits at the center of how organizations operationalize LLMs across research, product, and production workloads.

About the Role

OpenRouter routes almost a billion requests and more than 20 trillion tokens a day, across 80+ providers and thousands of endpoints. Every one of those providers can degrade, rate-limit, change behavior, or go down without warning. Our customers count on us to absorb that chaos so their apps never notice.

We're hiring our first AI Inference SRE to own the operational health of our provider supply. You'll make sure every endpoint we route to is fast, correct, and available, and that we detect and route around problems before customers do. You'll sit on the Provider Operations team, reporting to the Provider Operations Manager.

What You'll Do
  • Provider health and observability. Build and own monitoring for every provider and endpoint: latency, throughput, error rates, uptime, and output correctness. Set SLOs per provider tier and alert on them.

  • Detection and failover. Improve how quickly we detect degraded endpoints, and work with the routing team so traffic shifts away from them automatically.

  • Incident response. Own on-call for provider incidents: triage, mitigate, communicate with providers, run postmortems, and drive follow-ups to closure.

  • Provider accountability. Turn telemetry into scorecards and SLO reporting that providers act on, and be the technical escalation point when a provider's endpoint is misbehaving.

  • Quality regression detection. Build continuous canaries and evals that catch silent regressions (quantization changes, broken tool calling, truncated streams, pricing or usage-reporting mismatches), not just outright downtime.

  • Automate the toil. Replace manual provider-ops work (disabling endpoints, capacity changes, deprecations, rate-limit tuning) with safe, auditable tooling.

  • Capacity and launch readiness. Build tooling to load-test endpoints before big launches so day-zero traffic doesn't take them down.

About You
  • 4+ years in SRE, production engineering, or infrastructure roles running high-traffic, customer-facing systems.

  • Strong with observability tooling and practice: metrics, tracing, logs, SLOs/error budgets, alerting that is always actionable.

  • Capable software engineer who prefers writing tools over executing runbooks. TypeScript and/or Python.

  • Experienced with distributed systems failure modes: timeouts, retries, backpressure, partial outages, noisy neighbors.

  • Calm, clear incident commander who communicates well with external partners under pressure.

  • Understands, or is eager to learn deeply, how LLM inference is served: streaming, tool calling, prompt caching, throughput/latency tradeoffs, and how provider APIs differ.

Nice to Have
  • Experience at an inference provider, model lab, GPU cloud, or API gateway/CDN company.

  • Experience with our stack: TypeScript, Cloudflare Workers, Postgres, ClickHouse, GCP, Vercel.

  • Background in routing, load balancing, or traffic management systems.

  • Experience with evals or synthetic monitoring for ML systems.

Similar Jobs

An Hour Ago
Remote or Hybrid
245K-335K Annually
Senior level
245K-335K Annually
Senior level
Fintech • Machine Learning • Payments • Software • Financial Services
Leads enterprise AI engineering strategy and multi-team delivery of scalable, responsible AI systems. Oversees foundation model training, LLM inference, similarity search, guardrails, evaluation, governance, observability, and production optimization. Establishes responsible AI standards, makes technology decisions, develops long-term platform roadmaps, partners with research and risk teams, and attracts and mentors engineering talent.
Top Skills: AWSAws UltraclustersAzureC#C++CudaGoGCPHugging FaceJavaPythonPyTorchVectordbs
An Hour Ago
Remote or Hybrid
245K-335K Annually
Expert/Leader
245K-335K Annually
Expert/Leader
Fintech • Machine Learning • Payments • Software • Financial Services
Leads data science for consumer and developer experiences, partnering with engineers and product managers to deliver customer-focused products. Builds, evaluates, validates, and deploys machine learning models using large-scale numerical and textual data. Applies statistical modeling, A/B testing, clustering, classification, sentiment analysis, time series, and deep learning while translating technical insights into business outcomes. The role also includes team leadership, talent development, and evaluating emerging AI and cloud technologies.
Top Skills: SparkAWSCondaGenerative AiH2OMachine LearningPythonRRelational DatabasesScala
An Hour Ago
In-Office or Remote
292K-439K Annually
Senior level
292K-439K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Leads population health strategy, clinical innovation, care-model design, and evidence-based program development. Uses clinical, claims, and utilization data to identify intervention opportunities and evaluates products, partnerships, clinical guidelines, and care programs. Represents clinical perspectives with providers, health systems, clients, and executives while developing clinical talent. Requires an active medical license, board certification, clinical leadership, clinical practice, population health experience, strong analytics, and communication skills.

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account