NVIDIA

Senior Site Reliability Engineer, BCM - DGX Cloud

Posted 18 Hours Ago

In-Office or Remote

2 Locations

168K-334K Annually

Senior level

In-Office or Remote

2 Locations

168K-334K Annually

Senior level

The Senior Site Reliability Engineer will manage deployments, operations, and incident handling for large-scale AI GPU platforms while ensuring high performance and resilience in configurations.

The summary above was generated by AI

NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for over 25 years. It’s a unique legacy of innovation fueled by great technology—and dynamic people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent.

NVIDIANS immerse themselves in a diverse, supportive environment that encourages everyone to do their best work. Join the team and see how you can make a lasting impact on the world. NVIDIA Base Command Manager powers thousands of clusters worldwide, varying from a few to several thousands of nodes, and streamlines cluster provisioning, workload management, and infrastructure monitoring. It provides all the tools you need to deploy and run an AI data center. We take great pride in providing excellent, comprehensive support to our customers! Sr Site Reliability Engineer in this role will significantly impact and contribute to the overall success of both external customers running their clusters with NVIDIA solutions AND internal clusters used for research, operations, and next-generation projects.

What you’ll be doing:

Contributing to deployments and daily operations of large scale next-generation GPU platforms
Handling incidents in GPU clusters, bridging the gap between cluster operations and development
Designing and implementing small features in the Base Command Manager product to become intimately familiar with the workings of the product
Validating complex cluster configurations including Slurm and Kubernetes orchestrators for performance, scalability and resilience, ensuring they meet real-world customer scenarios.

What we need to see:

Bachelor's Degree or equivalent experience in Computer Science or related field.
8+ years of experience in site reliability engineering and/or software development roles.
Fluency in Python
In-depth knowledge of Linux and networking

Ways to stand out from the crowd:

Experience with C++, high-performance computing, Kubernetes and/or system administration would be an asset
Previous experience as a system admin running BCM/Bright Cluster Manager/Base Command Manager clusters is a definite plus.
Proficiency with cluster networking including InfiniBand and Spectrum-X

NVIDIA is widely considered one of the world's most desirable employers in technology. We have some of the world's most forward-thinking and passionate people working for us. If you're creative and autonomous, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 270,250 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until October 31, 2025.NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Top Skills

C++

Kubernetes

Linux

Python

Similar Jobs

Imubit

Senior Site Reliability Engineer

23 Days Ago

In-Office or Remote

Houston, TX, USA

Mid level

Artificial Intelligence • Machine Learning • Energy

The Site Reliability Engineer designs and maintains cloud infrastructure at Imubit, optimizing deployment processes, managing incidents, and collaborating with teams to enhance system reliability and performance.

Top Skills: AnsibleAWSAws Secrets ManagerGCPGitGoGrafanaHashicorp VaultKubernetesNew RelicPostgresPrometheusPythonSplunkTerraform

Overstory

Senior Site Reliability Engineer

14 Days Ago

Remote

Senior level

Software

As a Senior Site Reliability Engineer, you'll manage GCP infrastructure, improve incident processes, develop observability platforms, and advocate for reliability best practices.

Top Skills: GCPInfrastructure-As-CodeKubernetesUnix

CyberArk

Senior Site Reliability Engineer

9 Days Ago

In-Office or Remote

Newton, MA, USA

119K-165K Annually

Senior level

119K-165K Annually

Senior level

Security • Software

The Senior Site Reliability Engineer will manage AWS infrastructure, ensure SaaS reliability, automate platforms, and respond to production incidents.

Top Skills: AnsibleAWSCloudFormationCloudwatchDatadogGrafanaHelmKubernetesOpensearchPagerdutySaltTerraform

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories