Positron AI Logo

Positron AI

Technical Product Manager, AI Inference & Software

Posted 6 Hours Ago
Be an Early Applicant
Remote
Hiring Remotely in United States
200K-350K Annually
Expert/Leader
Remote
Hiring Remotely in United States
200K-350K Annually
Expert/Leader
Own technical product planning for AI inference software: define requirements across model coverage, numerics, inference engine and serving-stack features, roadmap planning, ecosystem engagement, documentation, lifecycle management, and automation using agentic AI. Translate model and system trends into engineering scope and partner with engineering and GTM teams to deliver a managed inference service.
The summary above was generated by AI
About Positron AI

Positron AI is building next-generation AI inference accelerators designed from the ground up for low-latency, high-throughput large language model inference. Our first-generation ASIC, Asimov, is a cutting-edge accelerator targeting frontier AI workloads, with additional generations already underway.

Role Overview

Positron AI is looking for a Technical Product Manager to own AI inference and software technical product planning end to end. In this role, you will be the person who translates where models and inference systems are heading into concrete, well-scoped requirements for our inference software stack, spanning model coverage, numerics, inference-engine features and modes, serving-stack capabilities, and our managed service.

This is a deeply technical planning role that sits at the intersection of engineering, go-to-market, and the broader inference ecosystem. You will track the model frontier as a discipline, convert that movement into engineering requests before it becomes a customer escalation, and serve as the connective tissue between our engineering organization, our GTM teams, and our ecosystem partners. You will also be expected to use agentic AI daily as a core part of how the planning function operates.

Key Responsibilities

Product Requirements and Roadmap

  • Serve as the leader at Positron for all aspects of AI inference and software technical product planning.
  • Write requirements for the inference software stack spanning model coverage, numerics, inference-engine features and modes, serving-stack features and modes, and managed-service capabilities.
  • Partner with engineering and GTM to build and communicate a clear, defensible roadmap.
  • Create scope and feasibility frameworks that convert model and inference-system innovations into tangible engineering requests.

Market and Frontier Tracking

  • Keep planning ahead of where models and inference systems are going, tracking the model frontier as a discipline and converting movement into requirements before it surfaces as a customer escalation.
  • Work closely with GTM teams to understand customer and market needs and feed them back into the roadmap.
  • Create competitive briefings covering inference providers, serving stacks, and adjacent hardware platforms.
  • Engage key ecosystem partners, including model labs, open-source runtimes, and serving and orchestration partners, to understand their technology roadmaps.

Lifecycle, Documentation, and Process

  • Define and manage the software product lifecycle, including versions and release trains, feature modes, model catalog, and deprecation policy.
  • Ensure our software products are well documented for both internal and customer-facing audiences.
  • Streamline and automate the product planning process using agentic AI.
Required Qualifications
  • 10+ years of experience across ML systems, inference infrastructure, or serving-stack engineering, with direct ownership of performance or architecture trade-offs.
  • Deep expertise in transformer internals at the operator level, including attention variants, MoE routing, KV-cache mechanics, and quantization formats along with their hardware implications.
  • Hands-on experience with production inference serving at scale, covering multi-tenancy, latency SLAs such as TTFT and TPOT, batching and scheduling, disaggregated serving, KV-cache management, and observability.
  • Working fluency in the open-source inference ecosystem, including vLLM and SGLang-class runtimes, kernels, model ingestion, and how models are released, quantized, and adopted in practice.
  • Strong performance analysis skills spanning models and systems, including utilization reasoning, tokens per dollar and tokens per watt arithmetic, and benchmark design, with the ability to build and defend the math personally.
  • Demonstrated experience in competitive landscaping and analysis of inference providers and serving stacks, gained at a model lab, an inference API provider, or an AI hardware company.
  • A proven ability to learn quickly and span the full stack, from model-architecture details up to fleet-scale serving systems, while staying current with the model and inference landscape.
  • Excellent communication and interpersonal skills, with comfort navigating uncertainty and driving a process of idea and decision socialization.
  • Confidence being the most technically grounded person in a GTM room and the most market-aware person in an engineering room.
  • A strong instinct for owning decision history, serving as the documented answer to "why did we choose X," including with executive leadership.
  • Daily, hands-on use of agentic AI in real technical work, building and running agent workflows for research, analysis, and requirements drafting, with the judgment to verify and own everything the agents produce.
Preferred Qualifications
  • Prior experience at an AI hardware or custom silicon company, with exposure to the realities of bringing a new accelerator platform to market.
  • Direct contribution to or close engagement with open-source inference runtimes or serving projects.
  • Experience defining and operating a managed inference service, including model catalog and deprecation policy.
  • A track record of building internal automation or agentic workflows that measurably improved a planning or research function.
Leveling & Scope

While this role is currently posted at a specific level, we are a growth-oriented organization and are open to hiring at a more senior level for the right candidate. Please note that this job description serves as a focused but generalized overview of the role; specific responsibilities and impact expectations will be tailored to the experience and seniority of the final hire.

Why Join Us?
  • You will shape the software roadmap for a purpose-built inference accelerator, working at the layer where model architecture, systems performance, and real customer workloads meet.
  • You will have unusually direct influence, defining what gets built across the inference stack and seeing it land in silicon-backed products that compete on performance per dollar and performance per watt.
Compensation & Benefits

The base salary range for this role is $200,000 – $350,000.


Please note that the figures provided represent the base salary range only and do not include other elements of our total compensation package, equity, or comprehensive benefits.


At Positron AI, we value the unique expertise each candidate brings. While the range above reflects our typical expectation for the position, we reserve the flexibility to exceed this range for candidates whose specialized skills, significant experience, or unique qualifications fall outside the standard scope of the role. Final offers are determined based on a variety of factors, including internal equity, and individual impact.

Benefits & Perks

We want you to do your best work and feel confident that you and your family are taken care of. That means comprehensive coverage, real time to rest, and support for your future.

Health and wellness

  • Fully company-paid medical, dental, and vision insurance for you and your dependents
  • Company-paid life and disability coverage, with voluntary options to add more
  • Supplemental hospital, critical illness, and accident coverage available

Time off and flexibility

  • Unlimited paid time off, we encourage everyone to truly unplug and recharge
  • 13 paid company holidays
  • Remote-first culture with a company-provided computer and home office setup

Compensation and future

  • Competitive salary and equity
  • 401(k) with company matching, eligible from day one
Visa Support

This position is open to candidates currently authorized to work in the U.S. We cannot provide new visa sponsorship for this role but are open to facilitating H-1B visa transfers for eligible candidates.


Equal Opportunity Employer. If you're excited about the role but don't meet every bullet, we'd still love to hear from you.


Similar Jobs

40 Seconds Ago
Remote
USA
190K-290K Annually
Expert/Leader
190K-290K Annually
Expert/Leader
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Lead enterprise systems and data architecture across business domains, define target-state architectures and integration patterns, rationalize shadow IT, guide sensitive-data placement in regulated multi-environment contexts, support major strategic programs, establish architecture governance and review practices, and coach engineering teams to raise architectural maturity.
Top Skills: APIsBatch Data MovementDatabricksEvent-Driven ArchitectureOracle ErpSAPSnowflake
40 Seconds Ago
Remote
USA
180K-270K Annually
Senior level
180K-270K Annually
Senior level
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Lead operational reliability and platform enablement for Databricks: build monitoring, CI/CD, deployment standards, compute and job policies, observability, runbooks, and governance to support secure, cost-aware, production data workloads across regulated environments. Mentor engineers and align platform with cloud/infrastructure and compliance requirements.
Top Skills: Ci/CdDatabricksDatabricks Asset BundlesDatabricks WorkflowsDelta LakeInfrastructure-As-CodeService PrincipalsUnity CatalogVersion Control (Git)
13 Minutes Ago
Remote or Hybrid
Boston, MA, USA
60K-80K Annually
Junior
60K-80K Annually
Junior
Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation
Perform end-to-end revenue recognition and close activities for subscription, services, and usage revenue. Execute contract reviews, fair value allocations, journal entries, reconciliations, and quarterly/annual ASC 606 assessments. Support automation and process-improvement projects, collaborate with global accounting and sales/renewals/order management teams, and assist external auditors with revenue testing and disclosures.
Top Skills: AIExcelMondayNetSuiteOnestreamSalesforce

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account