Stability AI Logo

Stability AI

Generative AI Inference Engineer

Sorry, this job was removed at 08:24 p.m. (EST) on Wednesday, Jul 01, 2026
Remote
Hiring Remotely in United States
Remote
Hiring Remotely in United States

Similar Jobs

14 Minutes Ago
Remote
United States
200K-230K Annually
Senior level
200K-230K Annually
Senior level
Computer Vision • Digital Media • Kids + Family • Mobile • Software • Sports
Lead architecture and hands-on development of the end-to-end video pipeline (live ingest, transcoding, storage, playback, and developer SDKs). Mentor engineers, collaborate across mobile/web/ML/platform teams, own monitoring and on-call support, and help scale video streaming and VoD capabilities for millions of streams and clips.
Top Skills: APIsAv1AWSC#C++ContainerizationDashDrmFfmpegGoH.264 (Avc)H.265 (Hevc)HlsKotlinMobile Broadcast FrameworksMpegMpeg-2 TsNode.jsPythonReactRtmpRustSdksSrtSwiftTranscodingTypescriptVideo PlayersVod StorageVp9Vvc
16 Minutes Ago
Remote
USA
115K-179K Annually
Mid level
115K-179K Annually
Mid level
Consumer Web • Healthtech • Professional Services • Social Impact • Software
Run and scale Talent Acquisition enablement and AI-adoption programs: design training and onboarding for recruiters and hiring managers, define success metrics, build dashboards, document processes, and partner cross-functionally to measure and iterate program impact.
16 Minutes Ago
Easy Apply
Remote
United States
Easy Apply
85K-110K Annually
Senior level
85K-110K Annually
Senior level
Edtech • Social Impact
Drive full-cycle technical recruiting for a rapidly scaling AI-native education nonprofit. Source, screen, and coordinate candidates using AI tools, partner with hiring managers to improve pipeline quality, and ensure a high-touch candidate experience during a 6-month contract.
Top Skills: Ai-Assisted Sourcing ToolsClaudeClaude CodeGoogle SuiteGemGreenhouseLinkedIn

Generative AI Inference Engineer

<Remote> 

About the role: 

We are seeking passionate Machine Learning Engineers to join our Inference team, focusing on the creative applications of generative AI models. The ideal candidate will have substantial experience developing and running inference for multi-modal models. A deep understanding of diffusion model architectures and familiarity with workflow tools like ComfyUI are a big plus. You will be expected to leverage and push the boundaries of state-of-the-art inference optimization techniques for multi-modal generative models. This role offers the opportunity to work alongside top researchers and engineers, utilizing cutting-edge high-performance computing resources to make a significant impact in the rapidly evolving field of generative AI.

Responsibilities:  

  • Lead efforts to drive the design, development of customer-facing multi modal ML inference systems.
  • Work with the Platform and Inference teams on building inference systems for the next generation of models, where you will work on areas such as optimization, model tuning and deployment.
  • Partner with leading cloud providers to deliver hosted Stability AI inference solutions.
  • Be a strategic thought partner for leaders across the organization on driving business impact through machine learning
  • Be part of the team to bring new Stability models and pipelines into existence
  • Prototype and productionize inference platform improvements and new features 

Qualifications:

  • 7+ years working on productionizing machine learning systems, including inference pipeline development
  • Expert level knowledge on writing and running python services at scale
  • 5+ years working on python scientific stack, pyTorch and at least one high-performance inference framework (e.g. Triton and TensorRT)
  • Deep understanding of Diffusion Architecture
  • Experience profiling and optimizing deep neural networks on Nvidia GPUs, using profiling tools such as NVIDIA Nsight
  • Experience with python-based image manipulation/encoding/decoding frameworks, such as OpenCV
  • Experience deploying to cloud orchestration systems such as Kubernetes and cloud providers such as AWS, GCP, and Azure
  • Experience with Docker
  • Ability to rapidly prototype solutions and iterate on them with tight product deadlines
  • Strong communication, collaboration, and documentation skills
  • Experience with the open-source ML ecosystem (HuggingFace, W&B, etc.)

Equal Employment Opportunity:

We are an equal opportunity employer and do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, veteran status, disability or other legally protected statuses.

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account