Cerence Inc. Logo

Cerence Inc.

Senior Principal AI Engineer

Posted 2 Months Ago
Remote
Hiring Remotely in USA
141K-226K Annually
Senior level
Remote
Hiring Remotely in USA
141K-226K Annually
Senior level
Design, build, and operate distributed GPU training infrastructure for large neural networks. Optimize multi-node/multi-GPU execution, diagnose compute/memory/network bottlenecks, improve stability and fault tolerance, and partner with research teams to productionize large-model training pipelines for reliable, high-throughput training at scale.
The summary above was generated by AI
A Moving Experience.

What You Will Work On  

  • Design and operate distributed training systems for large neural networks (autoregressive, diffusion, State Space Models etc.) across GPU clusters 
  • Optimise multi‑node, multi‑GPU execution to maximize throughput and utilization 
  • Diagnose & resolve bottlenecks across compute, memory, and network 
  • Improve training stability and fault tolerance at scale 
  • Partner with research and applied ML teams to productionize large‑model training pipelines  

 Core Responsibilities  

  • Distributed Training Infrastructure  
  • Build and optimize GPU cluster orchestration using: 
  • Slurm  
  • Kubernetes  
  • Ray  
  • RunAI  
  • Ensure efficient scheduling, isolation, and fairness across training workloads 
  • Communication & Networking  
  • Optimize and debug distributed communication using:  
  • NCCL  
  • RDMA  
  • InfiniBand  
  • NVLink  
  • Minimize networking bottlenecks that dominate end‑to‑end training time  

  

  • Training Frameworks  
  • Scale large-model training using:  
  • PyTorch Distributed  
  • Megatron‑LM  
  • DeepSpeed  
  • Own multi‑node launch configurations, failure recovery, and performance tuning  
  • Memory & Performance Optimization  
  • Apply advanced memory optimization techniques:  
  • Activation checkpointing  
  • ZeRO (Stage 1–3) and offload strategies  
  • Balance compute, memory, and communication to push model size and batch scale 

What Success Looks Like  

  •  GPU utilization consistently stays high (>80–90%)  
  • Training scales cleanly from single node to dozens or hundreds of GPUs 
  • Communication overhead is minimized and predictable  
  • Large training jobs run stably for days or weeks without failure  
  • New models can be trained faster, larger, and more reliably than before  

 Required Experience & Skills  

  •  Strongly Required  
  • Deep hands‑on experience with distributed systems or ML systems 
  • Experience running large‑scale workloads on GPU clusters  
  • Production experience with PyTorch distributed training  
  • Strong understanding of parallelism strategies (data, tensor, pipeline parallelism)  
  • Low‑level understanding of GPU communication and networking  
  •  Critical Technical Skills  
  • GPU orchestration: Slurm, Kubernetes, Ray, RunAI  
  • Communication libraries: NCCL, RDMA, InfiniBand, NVLink  
  • Training frameworks: PyTorch Distributed, Megatron‑LM, DeepSpeed 
  • Memory optimisation: activation checkpointing, ZeRO offload techniques  

  

  

Common Problems You’ll Be Solving  

  • Many teams fail at scale because:  
  • GPU utilization is low despite large clusters  
  • Networking and communication dominate training time  
  • Training jobs crash or become unstable at large scale  
  • You will be explicitly focused on eliminating these failure modes.  

 Ideal Background  

  •  This role is a strong fit for individuals who have worked as:  
  • ML Systems Engineer  
  • Distributed Systems Engineer  
  • AI Infrastructure Engineer  
  • HPC Engineer transitioning into ML  
  • Experience working with large language models or foundation models is a strong plus, but deep systems expertise is valued over pure model architecture experience.  

 Why This Role Matters  

 Without robust distributed training infrastructure, progress on large models stalls. This role directly enables:  

  • Larger models  
  • Faster iteration cycles  
  • More reliable research-to-production pipelines  

  

You will be building the foundation that makes large‑scale AI possible. 


What we offer 

We offer a generous compensation and benefits package (in addition to the base salary), including: 

·       Salary range $141,400.00 USD - $226,300.00 USD It is not typical for offers to be made at or near the top of the range. The actual salary will be determined based on experience and other job-related factors. 

·       Annual bonus opportunity 

·       Insurance coverage (medical, dental, vision, life, and disability) 

·       Paid time off 

·       Paid holidays 

·       Company contribution to the 401K

·       Equity awards for certain positions and levels 

·       Remote and/or hybrid work available depending on the position 

All compensation and benefits are subject to the terms and conditions of the underlying plans or programs, as applicable, and may be amended, terminated, or replaced from time to time. 


Cerence Inc. (Nasdaq: CRNC and www.cerence.com) is the global industry leader in creating unique, moving experiences for the automotive world. Spun out from Nuance in October 2019, Cerence is a new, independent company that has quickly gained traction as a leader in the automotive voice assistant space, working with all of the world’s leading automakers – from Ford and Fiat Chrysler to Daimler, Audi and BMW to Geely and SAIC – to transform how a car feels, responds and learns. Its track record is built on more than 20 years of industry experience and leadership and more than 500 million cars on the road today across more than 70 languages.  

 

As Cerence looks to the future and continues an ambitious growth agenda, we need someone to join the team and help build the future of voice and AI in cars. This is an exciting opportunity to join Cerence’s passionate, dedicated, global team and be a part of meaningful innovation in a rapidly growing industry. 

EQUAL OPPORTUNITY EMPLOYER

Cerence is firmly committed to Equal Employment Opportunity (EEO) and to compliance with all federal, state and local laws that prohibit employment discrimination on the basis of age, race, color, gender, gender identity, gender expression, sex, sex stereotyping, pregnancy, national origin, ancestry, religion, physical or mental disability, medical condition, marital status, citizenship status, sexual orientation, protected military or veteran status, genetic information and other protected classifications. Cerence Equal Employment Opportunity Policy Statement.

All prospective and current Employees need to remain vigilant when it comes to executing security policies in the workplace. This includes:

- Following workplace security protocols and training programs to familiarize with the ways to maintain a safe workplace.
- Following security procedures to report any suspicious activity.
- Having respect for corporate security procedures to allow those procedures to be effective.
- Adhering to company's compliance and regulations.
- Encouraging to follow a zero tolerance for workplace violence.

- Basic knowledge of information security and data privacy requirements (e.g., how to protect data & how to be handling this data).

- Demonstrative knowledge of information security through internal training programs.

HQ

Cerence Inc. Burlington, Massachusetts, USA Office

Burlington, MA, United States, 01803

Similar Jobs

3 Days Ago
Remote or Hybrid
212K-407K Annually
Senior level
212K-407K Annually
Senior level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Architect and engineer secure, scalable enterprise AI infrastructure for LLM training and inference, machine learning, agentic AI, and production platforms. Lead HPC, GPU, Kubernetes, cloud, networking, storage, reliability, observability, security, and automation strategies. Resolve complex performance and scalability challenges, establish engineering standards, evaluate emerging technologies, lead cross-functional initiatives, and mentor engineers in a regulated environment.
Top Skills: Accelerated ComputingAgentic AiCloud-Native PlatformsDistributed SystemsGpuHigh-Performance NetworkingHpcHybrid CloudInfrastructure AutomationKubernetesLlmsMachine LearningObservabilityPrivate CloudRack-Scale SystemsSreStorage
3 Days Ago
Remote or Hybrid
212K-407K Annually
Senior level
212K-407K Annually
Senior level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Leads the architecture and engineering of secure, scalable enterprise AI infrastructure. Designs platforms for LLM training and inference, machine learning, agentic AI, and production workloads using HPC, GPUs, Kubernetes, hybrid cloud, networking, storage, and automation. Establishes standards for reliability, security, observability, governance, and compliance; resolves complex performance challenges; evaluates emerging technologies; leads cross-functional initiatives; and mentors engineers.
Top Skills: Accelerated ComputingAgentic AiDistributed SystemsGpu ComputingHigh-Performance NetworkingHigh-Performance StorageHpcHybrid CloudInfrastructure AutomationKubernetesLlmsMachine LearningObservabilityPrivate CloudRack-Scale SystemsResponsible AiSre
3 Days Ago
Remote or Hybrid
Boston, MA, USA
212K-407K Annually
Senior level
212K-407K Annually
Senior level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Architect and lead enterprise AI infrastructure supporting LLM training and inference, machine learning, agentic AI, and production AI platforms. Define technical standards across HPC, GPUs, Kubernetes, hybrid cloud, networking, storage, reliability, observability, security, and governance. Solve complex performance and scalability challenges, evaluate emerging technologies, lead cross-functional initiatives, and mentor engineers in a highly regulated environment.
Top Skills: Accelerated ComputingAgentic AiCloud-Native PlatformsDistributed SystemsGpusHigh-Performance NetworkingHigh-Performance StorageHpcHybrid CloudInfrastructure AutomationKubernetesLlmsMachine LearningObservabilityPrivate CloudRack-Scale SystemsSre

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account