Positron AI Logo

Positron AI

Engineering Manager, Serving/API

Posted Yesterday
Be an Early Applicant
Remote
Hiring Remotely in United States
225K-350K Annually
Entry level
Remote
Hiring Remotely in United States
225K-350K Annually
Entry level
Lead and grow the Serving/API engineering team responsible for OpenAI-compatible APIs, tokenization, structured generation, tool calling, reasoning, speculative decoding, multimodal serving, and model integration. Set technical direction, establish release and conformance gates, improve latency and API fidelity, partner across compiler, hardware, and production teams, and translate customer needs into scalable engineering roadmaps and ownership structures.
The summary above was generated by AI

About Positron AI

Positron builds high-performance systems for production AI inference. Our Upstack Engineering organization turns accelerator hardware into reliable, scalable services that customers can deploy and operate with confidence. The team works across distributed systems, model-serving infrastructure, compilers and runtimes, accelerator platforms, and production operations.


Role Overview

Positron is seeking an Engineering Manager to lead Production Platform and Orchestration within our Upstack Engineering organization. This team owns the software and operating practices that provision, deploy, observe, upgrade, and reliably operate Positron systems in production. You will inherit a technically strong core team and help it grow into a durable organization capable of supporting a fleet that is expanding by several multiples.

This is a technical leadership role with real operational accountability. You will set direction, build the team, create clear ownership, and improve the systems and processes behind fleet orchestration, deployment lifecycle, observability, incident response, release automation, and production reliability. You will work closely with serving and API, model enablement, compiler and runtime, hardware, customer-facing, and data center partners.

The strongest candidate will combine systems depth with organizational judgment, moving comfortably between architecture, delivery, incidents, people development, and cross-functional planning. This description intentionally emphasizes outcomes and ownership over a fixed organizational chart. As the fleet and customer base grow, the function may develop dedicated groups for fleet orchestration and capacity, deployment lifecycle, reliability and observability, data center operations, customer production operations, and operational tooling.

About the role

We are seeking an Engineering Manager to lead Serving/API. This team owns the software between a customer's API request and the tokens Positron accelerators generate: what the model sees, and how its output becomes a correct, well-formed response. You will lead a strong team and grow it as Positron adds models, modalities, and serving capabilities.

The team owns:

  • The OpenAI-compatible serving layer. HTTP endpoints, request validation, SSE streaming, usage accounting and API-spec fidelity.
  • Tokenization and chat templates. HuggingFace tokenizers, chat-template rendering, and model-specific conversation formats such as OpenAI Harmony for GPT-OSS.
  • Tool calling and structured output. Tool-schema handling, tool-call parsing and serialization, and grammar-constrained decoding (llguidance) for JSON and function calls.
  • Reasoning and budgets. Reasoning-channel parsing, reasoning-effort controls, token limits and per-request budgets.
  • Speculative decoding. The request-side plumbing for draft models and draft trees, acceptance metrics, and the policies that decide when speculation pays off.
  • New modalities and frameworks. Vision-language model (VLM) input handling and integration paths with SGLang-style serving.

This is a technical leadership role with real product accountability. You will set direction, build the team, create clear ownership, and improve the systems behind API fidelity, tool calling, reasoning, speculative decoding, and new-model readiness. You will work closely with Production Platform and Orchestration, Compiler/Executor, hardware, and customer-facing partners.

What you will do
  • Lead, coach, and grow a team of engineers spanning API serving, tokenization, structured generation, and model integration.
  • Establish a clear technical and organizational roadmap for the serving layer: API surface, chat templates, tool calling, reasoning, budgets, and new-model support.
  • Ensure API fidelity: OpenAI-compatible behavior, streaming, usage accounting, and error handling that customers and their tooling can rely on.
  • Lead reasoning-model support and request budgets, including reasoning formats, reasoning-effort controls, token limits, and per-request accounting.
  • Partner with Compiler/Executor on speculative decoding: draft-model integration, request-side plumbing, acceptance metrics, and policies for when speculation pays off.
  • Lead the plan for vision-language model (VLM) support and SGLang interoperability, deciding what to adopt, what to build, and how it fits Positron's engine.
  • Keep serving-layer overhead off the critical path for time-to-first-token and streaming latency.
  • Keep the API layer independent of hardware topology as Positron adds platforms that run inference across multiple hosts.
  • Work with Production Platform and Orchestration to define clear ownership of the shared host-level load-balancing layer, including where request policies and protections live.
  • Build release gates (API conformance tests, tool-call and reasoning evals, regression suites) so new models and features ship with confidence.
  • Collaborate with Production Platform and Orchestration to turn new serving capabilities into supportable production endpoints.
  • Translate customer and business priorities into sequenced engineering work while protecting the team from reactive, unstructured requests.
  • Hire thoughtfully, develop emerging leaders, and create ownership boundaries that remain effective as the organization scales.
What success looks like

First 6 months

  • Build trust with the team and partner organizations; clarify ownership, decision rights, and the near-term hiring plan.
  • Baseline API conformance, tool-call accuracy, serving-layer latency, and the largest sources of customer-visible defects.
  • Establish release gates for API conformance and tool-calling accuracy that run on every release.
  • Produce an agreed roadmap covering speculative decoding, VLM support, and SGLang interoperability alongside near-term model and customer needs.

6 to 12 months

  • Grow the team and create durable ownership for API serving, structured generation, reasoning and budgets, and new-model integration.
  • Ship speculative decoding in production for at least one model, with measured acceptance rates and throughput gains.
  • Make API support for new models (chat templates, tool-call and reasoning formats) repeatable, with predictable turnaround.
  • Develop engineers and technical leads who can independently own major serving domains.

12 to 18 months

  • Support a substantially larger model catalog and feature surface without proportional growth in team effort.
  • Navigate transition to new hardware platform with expanded capabilities and serving requirements.
  • Demonstrate measurable improvement in API fidelity, tool-call accuracy, and time-to-first-token.
What we are looking for
  • Demonstrated success managing and growing engineering teams responsible for LLM inference, model serving, ML systems, or a closely related domain.
  • Systems depth in C++ and working proficiency in Python; comfort reviewing performance- and correctness-critical code.
  • A record of turning ambiguous product demands into a coherent roadmap, explicit ownership, and measurable engineering outcomes.
  • Ability to recruit, coach, and retain engineers across experience levels while maintaining a high technical bar.
  • Clear written and verbal communication, sound prioritization, and the ability to make tradeoffs visible to technical and executive stakeholders.
  • A hands-on leadership style: close enough to architecture and code to ask the right questions without becoming the team's bottleneck.
Especially valuable experience
  • Contributing to vLLM, SGLang, llguidance, XGrammar, Hugging Face tokenizers, or similar projects.
  • Strong technical judgment across LLM serving: tokenization, chat templates, sampling, streaming APIs, and the OpenAI Chat Completions and Responses APIs.
  • Hands-on understanding of tool calling, function-calling formats, and structured output, and of how they fail in practice.
  • Familiarity with open-source serving stacks such as vLLM, SGLang, or TensorRT-LLM, and the judgment to know when to adopt rather than build.
  • Serving vision-language models: image preprocessing, vision encoders, and multimodal token handling.
  • Supporting reasoning models and their output formats, such as OpenAI Harmony.
  • Serving on GPU, FPGA, ASIC, or other accelerators, especially memory-bandwidth-bound inference.
  • Customer-facing API products where engineering teams own compatibility, escalations, and release readiness.
Leadership profile

The strongest candidate will combine serving-systems depth with organizational judgment. They will be comfortable moving between API design, model integration, performance work, people development, and cross-functional planning. They will keep the team focused on correctness and compatibility as new models, modalities, and serving techniques arrive faster than any one team can absorb.

Role scope

This description intentionally emphasizes outcomes and ownership over a fixed organizational chart. As the model catalog and customer base grow, the function may develop dedicated groups for API and protocol compatibility, structured generation and tool calling, speculative decoding, and multimodal serving.

Why Join Us?
  • You will build the production platform that turns purpose-built inference silicon into services customers can depend on, with direct ownership of how a rapidly growing fleet is deployed, operated, and scaled.
  • You will shape both the technology and the organization from an early stage, defining the orchestration, reliability, and automation foundations that Positron will operate on for years to come.
Compensation & Benefits

The base salary range for this role is $225,000 – $350,000.

 

Please note that the figures provided represent the base salary range only and do not include other elements of our total compensation package, equity, or comprehensive benefits.

 

At Positron AI, we value the unique expertise each candidate brings. While the range above reflects our typical expectation for the position, we reserve the flexibility to exceed this range for candidates whose specialized skills, significant experience, or unique qualifications fall outside the standard scope of the role. Final offers are determined based on a variety of factors, including internal equity, and individual impact.

Benefits & Perks

We want you to do your best work and feel confident that you and your family are taken care of. That means comprehensive coverage, real time to rest, and support for your future.

Health and wellness

  • Fully company-paid medical, dental, and vision insurance for you and your dependents
  • Company-paid life and disability coverage, with voluntary options to add more
  • Supplemental hospital, critical illness, and accident coverage available

Time off and flexibility

  • Unlimited paid time off, we encourage everyone to truly unplug and recharge
  • 13 paid company holidays
  • Remote-first culture with a company-provided computer and home office setup

Compensation and future

  • Competitive salary and equity
  • 401(k) with company matching, eligible from day one
Visa Support

This position is open to candidates currently authorized to work in the U.S. We cannot provide new visa sponsorship for this role but are open to facilitating H-1B visa transfers for eligible candidates.

 

Equal Opportunity Employer. If you're excited about the role but don't meet every bullet, we'd still love to hear from you.


Similar Jobs

18 Minutes Ago
Remote or Hybrid
45K-85K Annually
Junior
45K-85K Annually
Junior
Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Handle inbound calls and warm leads, consult customers on insurance needs, recommend appropriate Property and Casualty coverage, and convert prospects into policyholders. The role includes paid licensing and training, commission opportunities, benefits, and remote work. Employees must work a weekday and weekend schedule, maintain a dedicated home office with high-speed wired internet, and obtain a Property and Casualty insurance license after hire.
Top Skills: Cable InternetDsl InternetFiber InternetPc
18 Minutes Ago
Remote or Hybrid
Boston, MA, USA
150K-185K Annually
Senior level
150K-185K Annually
Senior level
Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation
Owns infrastructure security across AWS, Azure, and GCP cloud environments. Defines security standards, IAM policies, secrets management, guardrails, and infrastructure-as-code controls. Conducts threat modeling, vulnerability assessments, and penetration testing; monitors alerts and leads incident response. Drives SOC 2 and ISO 27001 compliance, audit readiness, and AI-assisted threat detection. Partners with platform, engineering, operations, legal, and business stakeholders to remediate risks and strengthen cloud security.
Top Skills: AWSAws Secrets ManagerAzureCi/CdCloudFormationDarktraceGCPIamIso 27001OrcaSIEMSoc 2TerraformVaultVectraWiz
19 Minutes Ago
Remote
Virginia, USA
160K-250K Annually
Senior level
160K-250K Annually
Senior level
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Designs, develops, tests, validates, and deploys AI/ML models across generative AI, classification, prediction, computer vision, NLP, and reinforcement learning. Builds data and model-training pipelines, leverages cloud platforms, evaluates model performance, documents methodologies, collaborates with technical teams, and guides prototype-to-production transitions. Requires five years of relevant experience and an active TS/SCI clearance with polygraph.
Top Skills: Cloud ComputingComputer VisionFlaskGenerative AiGitKerasMachine LearningNatural Language ProcessingPythonPyTorchRReinforcement LearningTableauTensorFlow

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account