Vida Health Logo

Vida Health

Site Reliability Engineer III

Posted Yesterday
Remote
Hiring Remotely in United States
175K-185K Annually
Senior level
Remote
Hiring Remotely in United States
175K-185K Annually
Senior level
Serve as Vida’s first dedicated Site Reliability Engineer, modernizing Terraform, standardizing environments, improving CI/CD, scaling GCP and Kubernetes infrastructure, upgrading databases and runtimes, and strengthening monitoring, observability, security, and operational processes. The role includes building runbooks, on-call and escalation procedures, supporting enterprise launches, managing infrastructure costs, and evaluating multi-cluster Kubernetes architecture in a fully remote environment.
The summary above was generated by AI
ABOUT US
 
At Vida, we help people get better- and we're helping the healthcare system get better, too.
 
Vida is a virtual, personalized obesity care provider that uses evidence-based treatment to help patients manage obesity and related conditions like diabetes, high blood pressure, anxiety and depression. Vida's team of Obesity Medicine-Certified Physicians, Registered Dietitians, Expert Coaches and Licensed Therapists takes a whole-person approach to care, helping people lose weight, reduce stress and improve their overall health.
 
By combining advanced technology with top-notch healthcare providers, Vida is breaking down the barriers that have historically kept people from getting the best care. It's trusted by Fortune 100 companies, major national payers and large providers to enable their employees to live their healthiest lives.

Vida has been operating and growing for years, and our infrastructure reflects that. We run on GCP with a production GKE cluster hosting around 50 workloads, from Django applications to scheduled Airflow jobs. Our data layer includes Cloud SQL (MySQL and PostgreSQL), Redis, and Firestore. Our infrastructure is defined in two Terraform repositories, one for core GCP infrastructure and one for our data platform, and both have grown across many contributors over time. Until now, our infrastructure has been managed by backend engineers with deep infrastructure experience, and this role adds our first dedicated SRE to that group.

You'll be Vida's first dedicated Site Reliability Engineer. You'll join the Enablement Team, which owns the platform and tooling our Engineering Teams build on. You'll report to the Engineering Manager and work closely with the team's Lead Engineer, who sets technical direction and will mentor you. This is a fully remote role with no time zone restrictions.

You'll modernize, consolidate, and scale our infrastructure as Vida takes on a wave of new enterprise contracts starting January 1. You'll also help shape what SRE looks like at Vida going forward.

Repsonsibilities:

  • Consolidate our Terraform, which has grown into inconsistent patterns across our infrastructure and data repositories, into a clean, well-documented structure the whole team can work in. Establish conventions for state management, module structure, code review, and CI checks.
  • Normalize environments, improve build and deploy automation in GitHub Actions, and add drift detection and alerting.
  • Apply overdue patches and upgrades across our Cloud SQL databases and application runtimes.
  • Right-size compute and database workloads for growth, including connection pooling and scaling improvements for high-traffic services.
  • Evaluate our Kubernetes architecture as we grow, including whether and when to move to a multi-cluster setup.
  • Improve monitoring and observability in Datadog and Cloud Monitoring so we catch issues before they become incidents.
  • Design observability access for contractors and external partners that gives them the visibility they need while keeping protected health information out of view.
  • Retire legacy infrastructure and tooling that has been replaced but not yet decommissioned.
  • Build repeatable operational processes, including runbooks, an on-call rotation, and escalation documentation.
  • Support infrastructure readiness for Vida's January 1 enterprise launches.
  • Additional responsibilities as needed.

Qualifications:

  • Bachelor's degree at a minimum.
  • 5+ years of experience in SRE, DevOps, or infrastructure engineering, with real ownership of production systems.
  • Deep hands-on Terraform experience, including structuring modules and managing state across environments.
  • Strong working knowledge of GCP, including GKE, Cloud SQL (MySQL and PostgreSQL), IAM, networking and load balancing, and cost management.
  • Production Kubernetes experience, including autoscaling, resource management, and judgment about what belongs in the cluster versus outside it.
  • Hands-on experience building monitoring, alerting, and dashboards with tools like Datadog or Cloud Monitoring.
  • Proficiency in Python for tooling and automation.
  • Comfortable working across multiple teams and disciplines, and explaining infrastructure decisions to non-specialists.

Preferred:

  • Experience as an early or first SRE hire.
  • Experience refactoring or consolidating a large, organically grown Terraform codebase.
  • Experience improving observability from a less mature baseline.
  • Experience in a HIPAA-regulated or other compliance-driven environment.
  • CI/CD experience with GitHub Actions.
  • Experience running Django applications or Airflow in production on Kubernetes.
  • Experience designing or migrating to multi-cluster Kubernetes architectures.

Vida is proud to be an Equal Employment Opportunity and Affirmative Action employer.
 
Diversity is more than a commitment at Vida—it is the foundation of what we do. All qualified applicants will receive consideration for employment without regard to race, color, ancestry, religion, gender, gender identity or expression, sexual orientation, marital status, national origin, genetics, disability, age, or Veteran status. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law.
 
We seek to recruit, develop and retain the most talented people from a diverse candidate pool. We don’t just accept differences — we celebrate them, we support them, and we thrive on them for the benefit of our employees, our platform and those we serve. Vida is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures.
 
We do not accept unsolicited assistance from any headhunters or recruitment firms for any of our job openings. All resumes or profiles submitted by search firms to any employee at Vida in any form without a valid, signed search agreement in place for the specific position will be deemed the sole property of Vida. No fee will be paid in the event the candidate is hired by Vida as a result of the unsolicited referral.
 
**Vida is authorized to do business in many, but not all, states. If you are not located in or able to work from a state where Vida is registered, you will not be eligible for employment. Please speak with your recruiter to learn more about where Vida is registered.
 
Please note: Applicants must be authorized to work in the U.S. as Vida is unable to sponsor work visas for any position.
All Vida Employees must reside in/be able to work from the U.S.- international work is prohibited. Job postings at Vida will remain open through end of year, until filled.
 
#LI-remote

Similar Jobs

Yesterday
Remote
United States
125K-150K Annually
Senior level
125K-150K Annually
Senior level
Cloud • Information Technology
Own the reliability, scalability, and administration of production Vitess/MySQL and Cassandra databases. Design highly available architectures, replication, backup, recovery, security, and disaster recovery strategies. Lead incident response, on-call operations, monitoring, SLO development, root cause analysis, and production readiness reviews. Automate database and infrastructure operations using scripting, Kubernetes, Terraform, Ansible, and Jenkins. Establish runbooks, escalation procedures, training materials, and mentor junior database SREs while partnering across engineering and infrastructure teams.
Top Skills: AnsibleAWSAzureBashCassandraCatchpointDockerElkFirehydrantGCPGoGrafanaItilJenkinsKubectlKubernetesLinuxMySQLMysqlshNomadOssPrometheusPythonSQLTerraformVaultVitess
One Month Ago
In-Office or Remote
130K-153K Annually
Senior level
130K-153K Annually
Senior level
Consumer Web • Information Technology • Mobile • Other • Software • App development
Build, maintain, and automate onX's infrastructure platform and deployment pipeline using IaC. Manage Kubernetes/GKE, Terraform, GCP services, observability, and incident response. Improve performance, availability, cost, and developer path to production while participating in on-call rotations and collaborating on architecture.
Top Skills: AirflowBigQueryBigtableChecklyClaude CodeCloud RunCloud SqlCockroachdbGCPGkeGoogle Cloud MonitoringGoogle Cloud StorageGoogle ComposerIamKubernetesNoSQLOpentelemetryOpentofuPrometheusPub/SubRootlySQLTerraform
One Month Ago
Easy Apply
Remote or Hybrid
2 Locations
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills: AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account