Maximum of 25 job preferences reached.
Top Senior Site Reliability Engineer Jobs in Boston, MA
Fintech • Information Technology • Payments • Sharing Economy • Financial Services • Cryptocurrency
Leads reliability, scalability, performance, and security for large-scale cloud systems. Designs AWS infrastructure with Terraform, automates CI/CD and operational workflows, establishes SLOs, manages incident response and disaster recovery, develops monitoring and observability solutions, and builds internal tools. Partners with engineering teams on architecture and reliability practices, conducts code reviews, mentors SREs, and supports compliance, vulnerability management, and security integration.
Top Skills:
Agentic ApplicationsAmazon Api GatewayAmazon AuroraAmazon CloudfrontAmazon CloudwatchAmazon DynamodbAmazon EbsAmazon Ec2Amazon EcsAmazon EfsAmazon RdsAmazon Route 53Amazon S3Amazon VpcAWSAws FargateAws LambdaAws X-RayChaos EngineeringCi/CdDastDatadogDistributed SystemsDockerEvent-Driven SystemsGitlabGitopsGrafanaIamInfrastructure As CodeJavaKubernetesLlmsMicroservicesNew RelicNode.jsOwasp Top 10PythonSastServerless ArchitecturesSplunkTerraform
Automotive
Design, build, and operate a global observability platform across hybrid cloud and on-premises environments. Responsibilities include developing monitoring pipelines, infrastructure-as-code, reliability automation, SLI/SLO frameworks, performance optimization, incident response, root-cause analysis, and production troubleshooting. The role partners with engineering teams to improve system resilience, reduce toil, integrate AI/ML for anomaly detection, and establish observability best practices. It also provides technical mentorship and guidance.
Top Skills:
Ai/MlAmazon Web ServicesCC++Ci/CdDatadogDockerDynatraceElkGoGoogle Cloud PlatformJ2EeJavaKafkaKubernetesMicroservicesAzureNagiosNew RelicNoSQLOpentofuPrometheusPythonRestful ApisScalaSensuSplunkSpring BootSQLTcp/IpTerraform
Automotive
Develop and maintain global monitoring and observability platforms using Go, JavaScript, GCP, Kubernetes, OpenTelemetry, PostgreSQL, and Terraform. Improve reliability, scalability, performance, security, and disaster recovery for cloud services. Responsibilities include troubleshooting production systems, capacity planning, automation, on-call support, incident postmortems, code reviews, documentation, and vulnerability assessments.
Top Skills:
Document DatabasesDynatraceGoGoogle Cloud PlatformInfrastructure As CodeJavaScriptKubernetesOpentelemetryPostgresRelational DatabasesTerraform
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Design and implement reliable, scalable IT infrastructure; automate processes; monitor systems; resolve incidents; conduct performance and load testing; optimize cloud environments; maintain storage architecture; and lead continuous improvement. The role requires troubleshooting complex system issues, developing automation solutions, managing incidents, collaborating across teams, mentoring others, and supporting secure business operations across AWS, Google Cloud, and Microsoft Azure.
Top Skills:
AWSGoogle Cloud PlatformAzure
Reposted 15 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Information Technology • Security • Software • Cybersecurity
As a Staff Site Reliability Engineer, you'll oversee Zscaler production data center services, optimize code, and ensure cloud service availability and performance. Collaborate with cross-functional teams to improve processes and resolve escalated issues.
Top Skills:
BashDnsFirewallsGrafanaHTTPIcmpLoad BalancingNagiosOsi ModelPrometheusPythonTcp/Ip
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads architecture, modernization, optimization, and reliability initiatives for mainframe CICS, MQ, and z/OS Connect environments. Provides technical direction across development and operations teams, establishes governance and change processes, tunes performance using telemetry, resolves incidents, and develops modernization roadmaps. Collaborates with stakeholders and enterprise architects to deliver secure, scalable, high-availability solutions while evaluating automation, cloud integration, and AI technologies.
Top Skills:
AnsibleCicsCobolDevOpsIbm MqIbm Z/OsOpenshiftPythonRed Hat Ansible Automation PlatformZ/Os ConnectZlinux
14 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Security • Software • Cybersecurity • Automation
Provide technical direction for GitLab Dedicated, a managed single-tenant SaaS platform. Lead architecture and transformation across resilience, failover, tenant orchestration, change management, automation, and platform integrations. Identify systemic reliability and scalability risks, establish reusable platform patterns, strengthen service ownership, and guide cross-team technical decisions. Mentor senior engineers and advance engineering excellence across the organization.
Top Skills:
Cloud InfrastructureDevsecopsDistributed SystemsGoInfrastructure As CodeObservabilityPythonRuby
Reposted 23 Days AgoSaved
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Leads teams designing and deploying AI-driven enterprise and cloud security solutions. Responsibilities include developing machine learning systems, integrating data infrastructure and pipelines, performing advanced modeling and analysis, and building AI applications with Python, Java, and C++. The role involves client engagement, strategic problem-solving, stakeholder validation, coaching, process innovation, and operational excellence while addressing complex cybersecurity and privacy challenges.
Top Skills:
Ai SystemsAWSC++Cloud SecurityData EngineeringData PipelinesDatabricksGCPJavaMachine LearningAzurePythonScikit-LearnSnowflakeTensorFlow
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead long-term strategy and architecture for cloud and on‑prem platform infrastructure, driving Kubernetes and multi‑cloud reliability, IaC/GitOps automation, observability, SLO/SLI/error‑budget practices, incident leadership, AI‑augmented tooling adoption, and mentorship of senior engineers to improve platform resilience and developer experience.
Top Skills:
Amazon Elastic Kubernetes Service (Eks)AutoscalingAWSCapacity PlanningCi/CdGitopsGoGoogle Cloud PlatformGoogle Kubernetes Engine (Gke)Identity And Access ManagementInfrastructure As CodeKubernetesLinuxNetworkingObservabilityOperatorsPulumiPythonRke2StorageTerraform
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Build and maintain the Zero Touch Platform: develop Temporal workflows and Go services to automate cloud operations, own production systems (SLOs, on-call, incidents), debug cloud-native distributed systems, create CI/CD and IaC automation, and produce documentation to enable self-service across engineering teams.
Top Skills:
AksApmAWSAws CloudformationAzureC#Ci/CdEksGoJavaKubernetesLoggingMetricsPythonTemporalTerraform
Cloud • Software
Support Salesforce GovCloud reliability and high availability in a 24/7 operations environment. Responsibilities include incident response, root cause analysis, infrastructure administration, security and compliance alignment, automation, monitoring, preventative remediation, and workflow improvement. The role requires collaboration across engineering teams, participation in post-incident reviews, and support for AWS/C2S infrastructure and customer-facing services.
Top Skills:
Ai ToolsAWSAws CliAws SdksBambooBsdC2SChefGoJavaJenkinsKubernetesLinuxPuppetPythonRed Hat Enterprise LinuxSolarisSpinnakerTcp/IpUnix
Fintech • Payments • Financial Services
Leads the design and operation of highly available cloud systems, primarily on AWS, using Terraform, GitLab CI/CD, containers, and observability tools. Responsibilities include reliability engineering, incident response, disaster recovery, automation, monitoring, security integration, vulnerability management, and cost optimization. The role builds internal tools, guides architecture, mentors SRE engineers, documents operational processes, and partners with software engineering teams to improve production reliability and compliance.
Top Skills:
Agentic ApplicationsAmazon Api GatewayAmazon AuroraAmazon CloudfrontAmazon CloudwatchAmazon DynamodbAmazon EbsAmazon Ec2Amazon EcsAmazon EfsAmazon RdsAmazon Route 53Amazon S3Amazon VpcAWSAws FargateAws LambdaAws X-RayDastDatadogDockerGitlabGrafanaIamJavaKubernetesLlmsNew RelicNode.jsOwasp Top 10PythonSastSplunkTerraform
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3 • Infrastructure as a Service (IaaS)
Manage AWS and GCP cloud environments, scale infrastructure globally, shape technical architecture, and build reliable CI/CD pipelines. Automate security and compliance controls, improve infrastructure performance, and develop systems interacting with smart contracts across multiple blockchains. The role requires Terraform, shell scripting, GitHub Actions, Docker, production cloud operations, and observability experience, with Kubernetes, networking, and fintech compliance knowledge preferred.
Top Skills:
AWSCi/CdDockerFirewallsGCPGithub ActionsGkeGoHelmInfrastructure As CodeKubernetesLoad BalancersMtlsNode.jsPciShellSsl/TlsTerraformTypescriptVpcZero Trust
Cloud • Security • Software • Cybersecurity
Ensure reliability, scalability, and usability of network infrastructure for Akamai Connected Cloud. Define requirements and SLOs, build automation and CI/CD pipelines, collaborate with dev/QA to improve code and stability, troubleshoot complex network issues (on-call), and mentor teammates while driving architectural standards.
Top Skills:
AnsibleArgocdBashBirdChefFrrGithub ActionsGoGobgpJenkinsLinux NetworkingPuppetPythonSalt Stack
Cloud • Information Technology • Internet of Things • Professional Services • Software
Deploys, operates, and maintains resilient AWS and Kubernetes infrastructure and microservices for the Webex for Government environment. Monitors service health, troubleshoots production incidents, automates infrastructure provisioning, and supports 24x7 operations. The role strengthens security and FedRAMP compliance through IAM, logging, hardening, vulnerability remediation, and continuous monitoring. It also partners with engineering, security, and compliance teams to improve reliability, document procedures, create runbooks, and conduct incident and root-cause reviews.
Top Skills:
Amazon CloudwatchAWSAws CloudtrailBashCi/CdDockerElastic StackFedrampGitGitlabGoGrafanaJavaJenkinsKubernetesLinuxMicroservicesPrometheusPythonSplunkTerraform
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead the design, automation, and scaling of global compute infrastructure across data centers, cloud, and on-prem. Operate GitOps with Rancher Fleet/Flux/Helm, build self-healing tooling, own cluster autoscaling and capacity strategy, define SLOs using Datadog, and participate in on-call rotation while mentoring peers.
Top Skills:
AWSContainerdDatadogDockerFluxGCPGitopsGoHelmHpaInfrastructure As Code (Iac)KarpenterKedaKubernetesLinuxNutanixPythonRancher FleetVsphere
Consumer Web • Gaming • Mobile • News + Entertainment • Software
Leads reliability practices across Infrastructure Engineering by defining Service Level Objectives, Service Level Indicators, error budgets, observability standards, and reporting. Partners with engineering teams to connect infrastructure reliability to application and platform outcomes, analyzes distributed-system failure modes, builds Datadog dashboards and monitoring, and influences reliability strategy through technical guidance, mentoring, and executive communication.
Top Skills:
Amazon Web ServicesDatadogDistributed SystemsError BudgetsKubernetesObservabilityService Level IndicatorsService Level ObjectivesSre
Consumer Web • Gaming • Mobile • News + Entertainment • Software
Leads the strategy, architecture, and operation of Kubernetes-based cloud and on-premise infrastructure. Drives reliability engineering, Infrastructure as Code, GitOps, observability, automation, incident response, capacity planning, and cost optimization. Leads cross-functional platform initiatives, establishes SLOs and error budgets, mentors engineers, and advances AI-enabled engineering practices across the organization.
Top Skills:
Amazon Elastic Kubernetes ServiceAWSCi/CdDistributed SystemsGitopsGoGoogle Cloud PlatformGoogle Kubernetes EngineInfrastructure As CodeKubernetesLinuxObservabilityPulumiPythonRke2Terraform
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead reliability, scalability, and operational excellence of large-scale database platforms across cloud and on-prem. Build automation-first database infrastructure (Kubernetes operators, IaC, GitOps), drive monitoring/SLOs, incident leadership, performance and cost optimization, and partner with application teams on safe schema/migration practices. Mentor engineers and evaluate AI-assisted workflows to improve productivity and reliability.
Top Skills:
AerospikeArgocdAuroraClaudeCloud SqlCursorDatabase OperatorsEksFluxcdGithub CopilotGitopsGkeGoKubernetesMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Build and operate Kubernetes-based platform infrastructure, infrastructure-as-code workflows, datastore services, and self-service platform capabilities. Collaborate with application teams on platform consumption patterns, diagnose datastore performance issues, contribute to architectural standards and code reviews, and support distributed systems using cloud services, Kafka, observability platforms, and relational or NoSQL databases.
Top Skills:
AWSAzureCassandraDatadogGCPGitopsGrafanaKafkaKubernetesLlm Developer ToolsMicroservicesMySQLNew RelicNoSQLPostgresPulumiTerraform
29 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Security • Software • Cybersecurity • Automation
Build and operate reliable, scalable production infrastructure for GitLab’s user-facing services. Responsibilities include developing infrastructure automation and tooling, managing Kubernetes deployments, maintaining infrastructure as code, supporting CI/CD and GitOps, participating in on-call and incident response, improving observability and SLOs, troubleshooting production systems, and documenting operational practices. The role spans Intermediate through Senior Staff levels and requires strong software engineering, cloud, reliability, and asynchronous collaboration skills.
Top Skills:
AlertingAWSCi/CdGCPGitopsGoInfrastructure As CodeKubernetesLoggingMetricsRubySlisSlosTerraform
Social Media • Software
Design, implement, and operate infrastructure for high-scale production systems serving millions of users. Own reliability, observability, incident response, deployments, capacity planning, cost management, and operational excellence across bare-metal and cloud environments. Develop automation and performance tooling, improve production readiness, manage vendors, lead incident reviews, and mentor engineers on reliability and distributed-systems practices.
Top Skills:
Bare-Metal ServersCloud ServicesDatabasesDistributed SystemsGoKubernetesLinuxNetworkingObservability SystemsStorage Systems
Artificial Intelligence • Big Data • Healthtech • Software • Biotech
Owns resilience, performance, scalability, and operational cost improvements for a cloud-native platform. Designs and delivers full-stack solutions, builds monitoring and analysis capabilities, performs architectural and code reviews, and resolves complex production issues. The role requires pragmatic modernization judgment, AI-assisted development experience, customer or user accountability, and participation in a weekend on-call rotation.
Top Skills:
Ai-Assisted DevelopmentCloud-Native PlatformsDashboardsFull-Stack Web ApplicationsGrafanaLoggingMonitoring
Edtech • Professional Services
The Site Reliability Engineer will operate and improve critical, high-traffic systems across AWS and Azure. Responsibilities include infrastructure-as-code, automation, incident response, root-cause analysis, monitoring, alerting, disaster recovery, capacity planning, performance tuning, high availability, cloud security, CI/CD, and multi-cloud networking. The role also involves mentoring, documentation, regular DR testing, and on-call participation.
Top Skills:
ArmAWSAws CdkAws Direct ConnectAzureAzure DevopsAzure ExpressrouteBashBicepCi/CdCircleCICloudFormationDatadogEc2Github ActionsIamInfrastructure As CodeKubernetesOwaspPagerdutyPowershellPythonRdsS3SnykSonarqubeTerraformVpcVpn
Other
Build and operate AWS GovCloud environments within FedRAMP boundaries, including multi-account infrastructure, guardrails, logging, security tooling, and Terraform provisioning. Develop automated fleet release, patching, verification, CI/CD, compliance evidence, telemetry, SLO, and alerting systems. Support VM and appliance fleets at scale, participate in on-call, and engineer recurring incident causes away. The role also requires compliance-as-code, continuous monitoring, and test harnesses validating operational state.
Top Skills:
AlertingArtifact ManagementAws ConfigAws GovcloudCi/CdDashboardsFedrampFipsGrafanaKubernetesOpaOpentelemetryOscalSecurity ScanningSlosTelemetryTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Boston Companies Hiring Senior Site Reliability Engineers
See AllPopular Job Searches
All Filters
Total selected ()
No Results
No Results
















_1.png)






.jpg)



