Stratus Logo

Stratus

Senior/Lead Site Reliability Engineer

Reposted Yesterday
Remote
Hiring Remotely in United States
Senior level
Remote
Hiring Remotely in United States
Senior level
Owns production reliability for a B2B SaaS platform, including SLOs, error budgets, observability, alerting, incident response, recovery exercises, performance and capacity engineering, load testing, disaster recovery, and production readiness. The role writes automation and production code, leads incidents, coaches engineering teams, and improves reliability across Kubernetes, Azure, databases, messaging, and service infrastructure.
The summary above was generated by AI

Stratus, deriving from the Latin term meaning 'layer', offers an advanced set of MEP specific solutions that seamlessly layer across a contractor's entire workflow from design to fabrication to installation. Our team of seasoned industry experts, skilled technology leaders, innovators, and entrepreneurs understands that fabrication does not occur in isolation, and increasingly, it may not happen within your own fabrication shop. Through close relationships with our customers—who include some of the most innovative and largest MEP contractors—we have developed a suite of Stratus tools to digitize, automate, and optimize piping, plumbing, sheet metal, and electrical contracting. Stratus provides the software layer an MEP Contractor needs to optimize profits with true "Data Driven Contracting."


GENERAL DESCRIPTION:

The Senior / Lead Site Reliability Engineer is accountable for how Stratus behaves in production. Reporting to the Senior Director, Platform Engineering, this role brings genuine SRE discipline to a platform that MEP contractors run their fabrication shops on — where downtime does not mean a degraded experience, it means work stops on a job site. Stratus is a ~100-person, primarily remote, Series B software company growing quickly.

Unrelenting Reliability is one of our company values, and this is the role that operationalizes it. You will define what reliable means in numbers, instrument the system so we can see it, and close the loop from production signal back into engineering priority. This is an engineering role with a production mandate: you will write code, tune queries, build alerting, run load tests, and lead incidents — and you will be measured on customer-visible availability and latency, not on tickets closed.

The load-bearing problems on your plate are: (1) establishing service level objectives and an error budget the whole engineering organization operates against, with the measurement infrastructure to back them; (2) building the detection and response capability — high-signal alerting, clear on-call and escalation paths, per-system incident ownership, and well-exercised recovery paths — so that production problems are caught and resolved fast; and (3) building the performance and capacity engineering practice for our data and messaging layers, with headroom measured rather than assumed.

You will work across every engineering pod, with our platform and security functions, and with the customer-facing teams who see reliability problems first. The right candidate is comfortable being the person who says a number out loud and defends it, and is drawn to a place where the reliability practice is being built rather than maintained. Final level (Senior or Lead) will be determined by the depth and scope of experience demonstrated against the qualifications below, not years of experience alone.


*US Eastern time preferred


KEY RESPONSIBILITIES:

  • Define and own service level indicators, objectives, and error budgets for customer-facing services; build the measurement pipeline that makes them trustworthy and publish availability and latency against target on a regular cadence.
  • Build and own production observability: instrumentation standards, dashboards, and actionable alerting on our stack — AKS on Azure, Prometheus, Loki, Tempo, and Grafana, with Istio as the service mesh and Flux for GitOps — extending coverage across every environment.
  • Establish and run incident response practice — on-call rotation and paging paths, severity definitions, per-system incident ownership, escalation into and out of customer support, and blameless post-mortems in incident.io.
  • Set the recovery standard: define what rollback and recovery must demonstrate, and verify through regular exercise that every service meets it.
  • Lead performance and capacity engineering: query and index tuning, connection-pool sizing, capacity modeling against real load, and designing for graceful degradation under partial failure.
  • Work with the existing team that runs our k6 load, stress, spike, and soak testing against a production-like environment; help set per-endpoint latency thresholds tied to our SLOs and make the results a gate on delivery.
  • Drive production readiness: review new services and significant changes against readiness criteria covering instrumentation, alerting, failure modes, resource limits, and rollback before they ship.
  • Close the loop on remediation: track post-mortem action items through to verified production change.
  • Contribute to business continuity and disaster recovery planning, including backup and restore validation, failover design, and recovery objectives.
  • Partner with engineering pods to push reliability ownership outward — coach teams on instrumenting and operating their own services.
  • Write code: tooling, automation, instrumentation libraries, and fixes in the production codebase.

QUALIFICATIONS:

  • 6-8+ years of professional engineering experience, with 3-5+ years in a dedicated SRE or production engineering role at a B2B SaaS company.
  • Demonstrated ownership of SLOs and error budgets in production — you have defined them, measured them, argued about them with product leadership, and changed engineering behavior with them.
  • Deep, practical observability skills with Prometheus, Loki, Tempo, and Grafana or comparable: you have instrumented real systems and built alerting that pages on customer impact rather than on CPU.
  • Hands-on incident command experience at meaningful severity, including building on-call and escalation practice from the ground up.
  • Strong database performance skills — query profiling, index design, connection pooling, and diagnosing saturation under load. MongoDB experience strongly preferred; comparable document or relational depth acceptable.
  • Production Kubernetes experience (AKS preferred) sufficient to debug a live problem — pod scheduling, resource limits, networking, and service mesh behavior (Istio preferred).
  • Solid coding ability in at least one general-purpose language (Go, Python, C#, or TypeScript); willingness to work in a C#/.NET codebase.
  • Experience with load and performance testing tooling (k6, JMeter, Gatling, or comparable) and with turning results into engineering priority.
  • Production experience on Azure or AWS, with real understanding of the failure modes of managed services.
  • Fluency with AI-assisted engineering tooling and a track record of designing AI-leveraged workflows for your team — this is a graded expectation at every level at Stratus.
  • Excellent written communication: you write post-mortems, runbooks, and reliability reports that executives and engineers both read and act on.
  • Judgment and steadiness under pressure, and the credibility to tell engineering and product leadership something they do not want to hear.

NICE TO HAVE:

  • Experience establishing or maturing an SRE practice — defining the discipline, not inheriting it.
  • Experience with Sentry or comparable application error-monitoring platforms.
  • Experience operating event-driven and real-time systems — message brokers (Azure Service Bus, Kafka) and websocket or push layers (SignalR or comparable).
  • Experience operating MongoDB Atlas at production scale, including replica set topology and Atlas performance tooling.
  • Experience with durable workflow orchestration (Temporal or comparable).
  • Background in multi-region or multi-zone architecture and DR design against stated recovery objectives.
  • Experience with incident.io, PagerDuty, or comparable incident management platforms.
  • Familiarity with DORA metrics and with reliability work inside SOC 2 or NIST 800-171 scope.
  • Experience with legacy monolith reliability — improving the operational behavior of an ASP.NET or comparable application you cannot rewrite.
  • Domain interest in MEP, BIM, AEC, or construction technology.
  • Prior experience in a Series B / growth-stage company navigating the transition from product-market fit to scale.


E-VERIFY STATEMENT 
Stratus participates in E-Verify. After you join the team, we'll verify your eligibility to work in the U.S. by submitting information from your Form I-9 to the Social Security Administration and, if needed, the Department of Homeland Security. This process happens post-hire only - we never use E-Verify to pre-screen applicants. 
E-Verify Notice 
Right to Work Notice 

Similar Jobs

11 Days Ago
Remote or Hybrid
United States
168K-210K Annually
Senior level
168K-210K Annually
Senior level
Consumer Web • Gaming • Mobile • News + Entertainment • Software
Own database reliability, scalability, and operational excellence across cloud and on-premises environments. Build automated Kubernetes-based database platforms, infrastructure tooling, failover and backup systems, and self-healing capabilities. Lead observability, performance optimization, capacity planning, incident response, and safe migration practices. Partner with application teams, leverage AI for operational improvements, mentor engineers, influence architecture, and drive technical strategy for highly available database infrastructure.
Top Skills: AerospikeArgocdAuroraClaudeCloud SqlCursorEksFluxcdGithub CopilotGitopsGkeGoKubernetesKubernetes OperatorsMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform
24 Days Ago
In-Office or Remote
2 Locations
146K-264K Annually
Senior level
146K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Architect, develop, test, and distribute software, services, and infrastructure supporting Akamai’s cloud hypervisor platforms. Improve observability, automate infrastructure processes, troubleshoot complex distributed-system issues, mentor engineers, and participate in on-call service restoration. The role requires deep Linux, kernel, virtualization, ARM hardware, large-scale infrastructure, DevOps, and configuration-management expertise.
Top Skills: AnsibleArmDevOpsDistributed SystemsKvm/QemuLinuxLinux KernelNested VirtualizationNvidia GraceObservability InfrastructureSaltstack
One Month Ago
Remote or Hybrid
United States
168K-210K Annually
Senior level
168K-210K Annually
Senior level
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead reliability, scalability, and operational excellence of large-scale database platforms across cloud and on-prem. Build automation-first database infrastructure (Kubernetes operators, IaC, GitOps), drive monitoring/SLOs, incident leadership, performance and cost optimization, and partner with application teams on safe schema/migration practices. Mentor engineers and evaluate AI-assisted workflows to improve productivity and reliability.
Top Skills: AerospikeArgocdAuroraClaudeCloud SqlCursorDatabase OperatorsEksFluxcdGithub CopilotGitopsGkeGoKubernetesMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account