Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in Boston, MA
Reposted 3 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Information Technology • Security • Software • Cybersecurity
As a Staff Site Reliability Engineer, you'll oversee Zscaler production data center services, optimize code, and ensure cloud service availability and performance. Collaborate with cross-functional teams to improve processes and resolve escalated issues.
Top Skills:
BashDnsFirewallsGrafanaHTTPIcmpLoad BalancingNagiosOsi ModelPrometheusPythonTcp/Ip
Reposted 4 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills:
AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead the design, automation, and scaling of global compute infrastructure across data centers, cloud, and on-prem. Operate GitOps with Rancher Fleet/Flux/Helm, build self-healing tooling, own cluster autoscaling and capacity strategy, define SLOs using Datadog, and participate in on-call rotation while mentoring peers.
Top Skills:
AWSContainerdDatadogDockerFluxGCPGitopsGoHelmHpaInfrastructure As Code (Iac)KarpenterKedaKubernetesLinuxNutanixPythonRancher FleetVsphere
Fintech • Software
Lead SRE efforts for DFIN SaaS: ensure availability, performance, scalability, and automation. Implement monitoring, CI/CD, IaC, container orchestration, AI-enhanced observability, incident response, RCA, and runbook automation while collaborating across engineering teams.
Top Skills:
.NetAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC#Ci/CdCloud Ai ServicesContainersCosmosDatadogDynatraceEksFirewallHarnessIdera Sql Diagnostic ManagerInfrastructure As Code (Iac)JavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
Reposted 10 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Maintain and improve multi-cloud Kubernetes infrastructure, CI/CD (Argo Workflows/ArgoCD), observability, and networking. Build reliable continuous deployment tooling and onboarding flows, provide internal support, collaborate across Platform Engineering, contribute upstream (open-source/operators), and participate in a 24/7 on-call rotation to resolve deployment infrastructure issues.
Top Skills:
AlertingArgo WorkflowsArgocdAWSAzureCi/CdContainersDnsGCPGoKubernetesLinuxLoad BalancerObservabilityPythonService MeshTcp/IpTls
Reposted 4 Days AgoSaved
Software • Defense
Own reliability, scalability, and security for on-prem and AWS deployments. Build observability (Prometheus/Loki/Grafana/ELK), define SLOs/SLIs, lead incident response and postmortems, automate infrastructure (Terraform/Ansible), operate Kubernetes clusters, embed security/compliance controls, eliminate operational toil, and mentor teams.
Top Skills:
AlloyAnsibleAWSAws GovcloudBashCloudFormationDatadogElkGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubernetesLokiPrometheusPythonRmfStigsTerraform
Big Data • Cloud • Software • Database
Seeking a Site Reliability Engineer with expertise in networking and distributed systems for building secure multi-cloud infrastructure. Responsibilities include maintaining network architecture and ensuring reliable service-to-service communication, involving a 24/7 on-call rotation.
Top Skills:
AWSAzureBgpDnsGCPIpv6KubernetesLoad BalancingMtlsService MeshTcp/IpTlsVpcsVpns
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills:
AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills:
Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
Insurance
As a Site Reliability Engineer II, you will build, test, and maintain the technology infrastructure for Openly's insurance platform, focusing on automation, monitoring, incident response, and operational decisions.
Top Skills:
Aiven DebeziumArcgisBigQueryCircleCICloud FunctionsCloud RunCloudsqlComposer/AirflowDatadogFivetranGcp GcsGitGoGCPJupyter NotebooksKafkaKubernetesNuxtPostgresPub/SubPythonRSQLTailwindTerraformVuejsWebpack
Artificial Intelligence • Machine Learning
Lead development of AI-assisted reliability tooling, own incident response end-to-end, improve observability and SLO/SLI frameworks, scale single-tenant SaaS operations, mentor engineers, and reduce recurring operational toil through engineering and automation.
Top Skills:
Cloud PlatformsGoKubernetesLinuxLlm/Ai ToolingLogs And TracingObservability ToolingPythonSlo/Sli Frameworks
Artificial Intelligence • Cloud • Consumer Web • Productivity • Software • App development • Data Privacy
The role involves defining reliability strategies, leading initiatives across teams, enhancing monitoring and incident response, and mentoring engineers at Dropbox.
Top Skills:
Ai TechnologiesDebuggingDistributed SystemsIncident ResponseObservabilityReliability Risk ManagementSlasSlos
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Reposted 20 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills:
BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Reposted 13 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Information Technology • Security • Software • Cybersecurity
This internship role focuses on SRE skills, requiring collaboration and problem-solving in dynamic environments for Zscaler's Zero Trust Exchange team.
Top Skills:
AnsibleAws EcsKubernetesLinuxPythonTerraform
Reposted 14 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Security • Software • Cybersecurity • Automation
As a Cloud Cost Utilization SRE at GitLab, you'll manage cloud spending, improve tracking and optimization of cloud usage, and collaborate with finance and engineering teams to enhance cost efficiency across AWS and GCP.
Top Skills:
AnsibleAWSElkGCPGrafanaLokiMimirPrometheusTempoTerraform
Reposted 24 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills:
AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
The Senior Site Reliability Engineer ensures the reliability and performance of cloud-native Kubernetes platforms by building tools, facilitating self-service for engineers, and promoting best practices.
Top Skills:
ArgocdAWSAzureC#Ci/CdGitGoJavaKubernetesPulumiPythonTerraform
Cloud • Information Technology • Internet of Things • Professional Services • Software
Design, build, and maintain automation and deployment tooling to improve reliability and scalability across global cloud environments. Troubleshoot distributed systems, develop CI/CD and testing frameworks, automate cluster and environment provisioning, and collaborate with engineering and product teams to reduce operational overhead and support large-scale platform growth.
Top Skills:
AnsibleGitlab CiLinuxRspecRuby
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Lead architecture and build a Kubernetes-based, GitOps-driven platform and self-service datastore offerings. Drive IaC, observability, and platform automation; partner with application teams to diagnose and optimize datastore and messaging performance at scale.
Top Skills:
AWSAzureCassandraDatadogGCPGitopsGrafanaKafkaKubernetesLlm/Agentic Ai ToolingMySQLNew RelicPostgresPulumiTerraform
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Build and maintain the Zero Touch Platform: develop Temporal workflows and Go services to automate cloud operations, own production systems (SLOs, on-call, incidents), debug cloud-native distributed systems, create CI/CD and IaC automation, and produce documentation to enable self-service across engineering teams.
Top Skills:
AksApmAWSAws CloudformationAzureC#Ci/CdEksGoJavaKubernetesLoggingMetricsPythonTemporalTerraform
2 Days AgoSaved
Financial Services
Ensure availability, performance, and resiliency of 40+ mission-critical trade-processing applications. Lead incident response and RCA, support DR and failover, drive automation, deployments, monitoring, observability, and collaborate with development, infra, cloud, security, and operations teams to reduce operational risk.
Top Skills:
AutosysAWSDb2DockerGrafanaIbm Integration Bus (Iib)Ibm MqItilJ2EeJavaJclKafkaKubernetesLinuxMainframeOpenshiftOraclePythonServicenowSplunkSQLUnix
Information Technology
Lead AIOps and Site Reliability-focused engineer responsible for incident management, operational resilience, observability, and intelligent automation. Partner with NOC, Cloud, DevOps, and application teams to implement AIOps, reduce MTTD/MTTR, build AI-driven operational agents, and drive ServiceNow and Dynatrace-based monitoring and automation. Mentor teams and define SLAs, SLOs, and reliability strategies.
Top Skills:
AiopsAWSAzureDevOpsDynatraceEnterprise AutomationGCPItilItsmMlopsMonitoringObservabilityPredictive AnalyticsServicenowSreWorkflow Orchestration
Cloud • Security • Software • Cybersecurity
Lead reliability for a serverless AI inference platform: own observability and SLO/SLI frameworks, build automation and tooling, manage incidents and on-call, define deployment safety (canaries, rollbacks), influence architecture with product teams, and mentor other SREs.
Top Skills:
AutoscalingCi/CdContainer OrchestrationContainerizationGoGpu WorkloadsInfrastructure-As-CodeKubernetesModel ServingPythonResource Scheduling
Hardware • Quantum Computing
Lead integration, maintenance, and automation of heterogeneous hardware and software control systems for quantum computers. Manage lab and test infrastructure (HIL, servers, networking, K8s), implement CI/CD and IaC, maintain observability and artifact pipelines, support incident response and root cause analysis, and establish operational best practices.
Top Skills:
AnsibleBashCi/CdDebianDhcpDnsDockerElkGitGitlab CiGoGrafanaHardware-In-The-LoopJenkinsKubernetesPrometheusPythonRack Mount ServersRed HatRoutersSwitchesTcp/IpTerraformUbuntuVlanWindows
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Boston, MA Companies Hiring SRE Engineers
See AllPopular Boston, MA Engineering Job Searches
Engineering Jobs in Boston, MA
Software Engineer Jobs in Boston, MA
Android Developer Jobs in Boston, MA
C# Jobs in Boston, MA
C++ Jobs in Boston, MA
DevOps Jobs in Boston, MA
Front End Developer Jobs in Boston, MA
Golang Jobs in Boston, MA
Hardware Engineer Jobs in Boston, MA
iOS Developer Jobs in Boston, MA
Java Developer Jobs in Boston, MA
Javascript Jobs in Boston, MA
Linux Jobs in Boston, MA
Engineering Manager Jobs in Boston, MA
.NET Developer Jobs in Boston, MA
PHP Developer Jobs in Boston, MA
Python Jobs in Boston, MA
QA Jobs in Boston, MA
Ruby Jobs in Boston, MA
Salesforce Developer Jobs in Boston, MA
Scala Jobs in Boston, MA
Application Engineer Jobs in Boston, MA
Automation Engineer Jobs in Boston, MA
AWS Engineer Jobs in Boston, MA
Backend Engineer Jobs in Boston, MA
Cloud Engineer Jobs in Boston, MA
Controls Engineer Jobs in Boston, MA
CTO Jobs in Boston, MA
Design Engineer Jobs in Boston, MA
DevOps Engineer Jobs in Boston, MA
Director of Engineering Jobs in Boston, MA
Electrical Engineering Jobs in Boston, MA
Embedded Software Engineer Jobs in Boston, MA
Full-Stack Engineer Jobs in Boston, MA
Game Engineer Jobs in Boston, MA
Infrastructure Engineer Jobs in Boston, MA
Manufacturing Engineer Jobs in Boston, MA
Mechanical Design Engineer Jobs in Boston, MA
Mechanical Engineer Jobs in Boston, MA
Mechatronics Engineering Jobs in Boston, MA
Network Engineer Jobs in Boston, MA
Platform Engineer Jobs in Boston, MA
Principal Engineer Jobs in Boston, MA
Principal Software Engineer Jobs in Boston, MA
Process Engineer Jobs in Boston, MA
Product Engineer Jobs in Boston, MA
Project Engineer Jobs in Boston, MA
QA Engineer Jobs in Boston, MA
Robotics Engineer Jobs in Boston, MA
Security Engineer Jobs in Boston, MA
Software Engineering Manager Jobs in Boston, MA
Software Test Engineer Jobs in Boston, MA
Solutions Architect Jobs in Boston, MA
Solutions Engineer Jobs in Boston, MA
SRE Engineer Jobs in Boston, MA
Staff Software Engineer Jobs in Boston, MA
Systems Engineer Jobs in Boston, MA
All Filters
Total selected ()
No Results
No Results





.png)








.jpeg)














