Ophelia Logo

Ophelia

Senior Site Reliability Engineer (Systems Engineer III)

Posted 3 Hours Ago
Remote
Hiring Remotely in United States
140K-150K Annually
Senior level
Remote
Hiring Remotely in United States
140K-150K Annually
Senior level
Own the reliability, availability, performance, observability, and security of Ophelia’s GCP telehealth platform. Define SLOs, improve monitoring and incident response, expand Terraform infrastructure, strengthen CI/CD, manage disaster recovery, and address latency and Cloud Run performance. Lead on-call maturity, HIPAA/SOC 2 infrastructure controls, and reliability education for engineers. Use AI-assisted tools for automation and incident triage while partnering with technical and non-technical stakeholders.
The summary above was generated by AI
Are you looking for a role in a company that's solving one of the greatest challenges of our lifetime? Ophelia helps people end their opioid use and restore their quality of life with respect for their time and dignity. Our mission is to make evidence-based treatments for opioid use disorder (OUD) accessible to everyone... and we're looking to bring more people onto our team to help us achieve it.
 
Ophelia is a venture-backed, healthcare startup that helps individuals with OUD by providing FDA-approved medication and clinical care through a telehealth platform. Our approach is discreet, convenient, and affordable. We've been successfully operating in 16 states for almost six years and we're excited to continue our growth. We are a team of physicians, scientists, entrepreneurs, researchers and White House advisors, backed by leading technology and healthcare investors working to re-imagine and re-build OUD treatment in America.
About the Role

As a Senior Site Reliability Engineer (SE III) at Ophelia, you will be our dedicated reliability, availability, and performance hire, playing a key role in making the systems that support our mission of treating opioid use disorder through telehealth stable, observable, and fast. Our stack is TypeScript on Node, React, and Firebase (e.g., Firestore, Authentication, Hosting) running on Cloud Run in Google Cloud Platform (GCP). Your work will have a direct impact on patients, clinicians, and our ability to scale.

You will own and drive our system stability and availability initiative, which is already underway: Terraform-driven uptime checks feed our engineering key performance indicators (KPIs), we have an initial service level objective (SLO) plan, and we've recently revamped our on-call rotation and runbook to keep the on-call engineer focused on severity incidents (SEVs). You will take this from a good start to a mature practice: defining and operationalizing SLOs, building out logging, monitoring, and alerting grounded in sound measurement (e.g., percentile-based latency indicators rather than averages), moving incident response into PagerDuty with clear categorization and escalation paths, expanding our infrastructure as code (IaC) footprint, and working down known pain points such as long-latency endpoints and Cloud Run cold starts. You will also partner closely with our head of IT on IAM hardening, secrets management, audit logging, and HIPAA/SOC 2 infrastructure controls. You will be the directly responsible individual (DRI) for our mean time to recovery (MTTR) KPI and for the measurement and reporting of our availability KPIs.

As a senior engineer, you will be a force multiplier for our team of fullstack engineers who are focused on product work. That means leveling up the team's GCP and DevOps skills through documentation, runbooks, pairing, and code review, and bringing reliability recommendations to new features as they are designed, while primarily owning the infrastructure work yourself. This role is entirely reliability-focused for the first six months; after that, you may contribute to product work as needed, though product work will be at most about 25% of your time.

Ophelia's Technology team actively encourages and invests in AI-augmented ways of working: using tools like Claude and Gemini to accelerate infrastructure code, runbooks, log analysis, and incident triage, and staying current as the AI landscape evolves. Consistent with Ophelia's AI Position Statement, we treat AI as a force multiplier for our team, not a replacement for human judgment: it's there to help you move faster from alert to root cause and spend more time on the reliability and architecture decisions that actually require a person. You'll be expected to use AI fluently in your own workflow, to build it into our operational tooling, and to help shape how the larger technology team uses it responsibly and effectively as it evolves.

Together, we will help hard-to-reach individuals treat their opioid dependence. While direct experience in this treatment area is not mandatory, knowledge of the healthcare space, including understanding health outcomes, benchmarks, systems, and regulatory compliance (HIPAA), is highly beneficial.

Responsibilities
  • Own reliability, availability, and performance of our GCP platform. Define and operationalize SLOs and service level indicators (SLIs) with the team, and drive the work that moves them (e.g., endpoint latency, Cloud Run cold starts, time to recovery), and own our disaster recovery posture (e.g., Firestore backup and restore testing, recovery point and recovery time objectives, and multi-region resilience).
  • Build observability and incident response. Extend our GCP Cloud Operations setup (log-based metrics, alerting policies, dashboards), move alerting from Slack into PagerDuty, participate in our on-call rotation alongside the rest of the team, and mature that process with clear incident categorization, escalation paths, and an AI-assisted level-one triage tool.
  • Grow our infrastructure as code and delivery pipeline. Expand our Terraform footprint beyond observability metrics and ephemeral pull request (PR) environments to the rest of our GCP configuration, and improve our GitHub Actions CI/CD pipeline so that deploying four times a day gets safer and faster, not riskier.
  • Harden security and compliance controls in partnership with our head of IT, who is the DRI for IT security and our HIPAA/SOC 2 infrastructure controls.
  • Multiply the team. Level up our engineers' GCP and DevOps skills through presentations, docs, runbooks, pairing, and code review; bring reliability recommendations to new features early; share reusable AI skills and automation across the team; and set the bar for operational excellence by embodying Ophelia's core values.
  • Take ownership of your work in our remote, growth-stage environment, driving projects from inception to impact with autonomy and proactive communication, adapting as our priorities evolve, and staying current with industry practice (including AI tooling) in reliability and cloud infrastructure.
Skills & Competencies
  • 5+ years experience in engineering, including 3+ in a cloud infrastructure, DevOps, or SRE-focused role, ideally with a portion at a growth-stage company with bespoke systems.
  • Expert-level, hands-on GCP fluency, including serverless compute (e.g., Cloud Run), Cloud Operations (Logging, Monitoring, alerting), IAM, and Firebase/Firestore. Terraform fluency is strongly preferred.
  • Expert in at least one language used for infrastructure automation (e.g., Python) and strong in shell scripting. You should be comfortable reading and debugging TypeScript/Node, as that's our entire application codebase; SQL for BigQuery and Log Analytics is a plus.
  • Production operations experience, including participating in on-call, leading incident response, writing postmortems, and defining SLOs in keeping with Google SRE practice, and the judgment to adapt those practices to a small team's capacity rather than importing large-company process wholesale.
  • Quantitative fluency applied to production systems, e.g., percentiles and tail latency rather than averages, SLO and error-budget math, burn-rate alerting, and reading time series and histograms well enough to tell signal from noise.
  • Senior-level ownership and communication, with the ability to own complex, ambiguous projects end-to-end with technical and non-technical stakeholders, clear and proactive communication, and a habit of writing things down (runbooks, postmortems, design docs) and pairing so that knowledge doesn't live only with you.
  • Daily, hands-on use of AI tools in your own engineering work (e.g., Claude Code for Terraform, runbooks, and log analysis), and experience building automation with them rather than only using them.
Preferred
  • Passionate about our mission to make evidence-based addiction treatment universally accessible.
  • Background in a regulated industry, ideally healthcare - telehealth is a big plus - with hands-on experience implementing HIPAA or SOC 2 infrastructure controls.
Our Benefits Include
  • Competitive medical, vision, and health insurance (many plans are fully covered for the employee!)
  • Start with 20 days (4 weeks) of PTO, increasing to 5 weeks after 2 years and 6 weeks after 5 years of tenure
  • 10 company holidays
  • Work From Home Stipend
  • 401k Contribution Platform
  • Additional benefits offered through our benefits provider such as life insurance, short and long term disability, financial wellness, virtual primary care, among others!

#LI-Remote

Ophelia Compensation Overview
  • We set compensation based on the level and skills required for the role. We value pay transparency and equity, and are committed to fair pay. In order to prevent pay disparities and reduce time spent in negotiations, we take a “first and best” offer approach: this means we’re not holding any compensation back from our candidates, and you can feel confident that our pay is fair and does not vary based on the strength of someone’s negotiation skills.
  • Compensation is dynamic at Ophelia: as long as the company performs well and meets our targets, there will be opportunities for increased compensation annually. We’re happy to discuss this approach and our bands if you have questions during the interview process.
Compensation Range
$140,000—$150,000 USD

Interested in learning more about Ophelia and this role? Apply to work with us! 

A Note Before Applying

We believe a good hiring process is really just about getting to know each other. That's why we care about getting to know the real you, not a polished, AI-generated version of you. Every application is read by a real person who's genuinely curious about your experiences, your voice, and why this role speaks to you.

So as you put together your application, we'd love for it to sound like you. It's the best way for us to get to know each other, and the best way for you to find out if we're the right fit.

Similar Jobs

A Minute Ago
Easy Apply
Remote or Hybrid
Easy Apply
152K-190K Annually
Junior
152K-190K Annually
Junior
Cloud • Information Technology • Security • Software • Cybersecurity
Design, train, fine-tune, optimize, and deploy large-scale machine learning systems for cloud security use cases. Build end-to-end ML pipelines, develop transformer and embedding models, productionize open-weight language models, and optimize inference for latency, cost, and quality. Architect resilient ML services across AWS and GCP using cloud-native microservices while collaborating with engineering teams on AI strategy and solving complex problems involving massive datasets.
Top Skills: AWSDeep LearningGCPHugging FaceJaxLarge Language ModelsLoraMicroservicesOnnx RuntimePeftPythonPyTorchQloraTensorFlowTensorrt-LlmTransformer ModelsVllm
2 Minutes Ago
Remote
USA
152K-175K Annually
Senior level
152K-175K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Secure Runpod’s multitenant GPU cloud infrastructure across bare-metal and virtualized environments. Responsibilities include designing workload and network isolation, hardening Linux kernels and container platforms, testing hypervisor and hardware security, developing low-level controls in C, Go, or Rust, securing GPU and PCIe interactions, and responding to infrastructure security incidents with forensic capabilities for ephemeral containers.
Top Skills: ApparmorBare-Metal Cloud InfrastructureCC.V.E.SCgroupsContainerdDockerEbpfGoGpu ArchitectureKubernetesKvmLinuxLinux KernelNamespacesPciePythonQemuRustSelinuxVirtualized Networking
4 Minutes Ago
Remote or Hybrid
153K-261K Annually
Entry level
153K-261K Annually
Entry level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Leads business development for BAE Systems’ Advanced Mission Solutions portfolio, shaping domestic and international growth strategies, identifying and qualifying opportunities, and developing customer engagement and campaign plans. Builds relationships with defense customers, industry partners, and internal teams; translates mission needs into product and technical roadmaps; supports capture, pricing, market analysis, compliance, and business-winning activities. Requires F-16 customer relationships, fighter aircraft knowledge, government contracting expertise, and approximately 50% international travel.
Top Skills: C4IsrDirect Commercial Sales (Dcs)F-16F-35Foreign Military Sales (Fms)

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account