Maximum of 25 job preferences reached.
Top Senior Site Reliability Engineer Jobs in Boston, MA
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead long-term strategy and architecture for cloud and on‑prem platform infrastructure, driving Kubernetes and multi‑cloud reliability, IaC/GitOps automation, observability, SLO/SLI/error‑budget practices, incident leadership, AI‑augmented tooling adoption, and mentorship of senior engineers to improve platform resilience and developer experience.
Top Skills:
Amazon Elastic Kubernetes Service (Eks)AutoscalingAWSCapacity PlanningCi/CdGitopsGoGoogle Cloud PlatformGoogle Kubernetes Engine (Gke)Identity And Access ManagementInfrastructure As CodeKubernetesLinuxNetworkingObservabilityOperatorsPulumiPythonRke2StorageTerraform
5 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Security • Software • Cybersecurity • Automation
Build and operate reliable, scalable production infrastructure for GitLab’s user-facing services. Responsibilities include developing infrastructure automation and tooling, managing Kubernetes deployments, maintaining infrastructure as code, supporting CI/CD and GitOps, participating in on-call and incident response, improving observability and SLOs, troubleshooting production systems, and documenting operational practices. The role spans Intermediate through Senior Staff levels and requires strong software engineering, cloud, reliability, and asynchronous collaboration skills.
Top Skills:
AlertingAWSCi/CdGCPGitopsGoInfrastructure As CodeKubernetesLoggingMetricsRubySlisSlosTerraform
Artificial Intelligence • Big Data • Healthtech • Software • Biotech
Owns resilience, performance, scalability, and operational cost improvements for a cloud-native platform. Designs and delivers full-stack solutions, builds monitoring and analysis capabilities, performs architectural and code reviews, and resolves complex production issues. The role requires pragmatic modernization judgment, AI-assisted development experience, customer or user accountability, and participation in a weekend on-call rotation.
Top Skills:
Ai-Assisted DevelopmentCloud-Native PlatformsDashboardsFull-Stack Web ApplicationsGrafanaLoggingMonitoring
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Leads teams designing and deploying AI-driven enterprise and cloud security solutions. Responsibilities include developing machine learning systems, integrating data infrastructure and pipelines, performing advanced modeling and analysis, and building AI applications with Python, Java, and C++. The role involves client engagement, strategic problem-solving, stakeholder validation, coaching, process innovation, and operational excellence while addressing complex cybersecurity and privacy challenges.
Top Skills:
Ai SystemsAWSC++Cloud SecurityData EngineeringData PipelinesDatabricksGCPJavaMachine LearningAzurePythonScikit-LearnSnowflakeTensorFlow
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Leads the design and development of AI-driven enterprise and cloud security solutions. Manages teams delivering data, analytics, and machine learning engineering projects; oversees AI model implementation, data infrastructure, pipelines, and platform deployments. Uses Python and C++ for algorithm development, analyzes data, improves data quality, ensures technical compliance, mentors staff, manages client expectations, and drives innovation across cybersecurity and privacy initiatives.
Top Skills:
AWSC++DatabricksGCPAzurePythonSnowflake
Reposted 22 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Leads reliability standards across Infrastructure Engineering by developing Service Level Objectives, Service Level Indicators, error budgets, observability reporting, and reliability measurement practices. Partners with engineering teams to connect infrastructure performance to application and customer experiences, analyzes distributed-system failure modes, and influences teams through technical guidance, mentoring, and communication with senior leadership.
Top Skills:
Amazon Web ServicesDatadogDistributed SystemsKubernetesObservability PlatformsOn-Premise Infrastructure
Cloud • Security • Software • Cybersecurity
Deploys and operates scalable, highly available cloud systems; improves application and network security, stability, speed, and capacity; automates cloud deployments; monitors and troubleshoots services to meet SLAs; analyzes logs and events; manages large-scale cloud infrastructure; and resolves infrastructure issues.
Top Skills:
Amazon Ec2Amazon EventbridgeAmazon Route 53Amazon S3Amazon Web ServicesAnsibleAws LambdaBashCi/CdDockerElastic Load BalancingJenkinsJinjaKubernetesPackerPulumiPython
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Build and maintain the Zero Touch Platform: develop Temporal workflows and Go services to automate cloud operations, own production systems (SLOs, on-call, incidents), debug cloud-native distributed systems, create CI/CD and IaC automation, and produce documentation to enable self-service across engineering teams.
Top Skills:
AksApmAWSAws CloudformationAzureC#Ci/CdEksGoJavaKubernetesLoggingMetricsPythonTemporalTerraform
Reposted One Month AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills:
AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
Reposted One Month AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Maintain and improve multi-cloud Kubernetes infrastructure, CI/CD (Argo Workflows/ArgoCD), observability, and networking. Build reliable continuous deployment tooling and onboarding flows, provide internal support, collaborate across Platform Engineering, contribute upstream (open-source/operators), and participate in a 24/7 on-call rotation to resolve deployment infrastructure issues.
Top Skills:
AlertingArgo WorkflowsArgocdAWSAzureCi/CdContainersDnsGCPGoKubernetesLinuxLoad BalancerObservabilityPythonService MeshTcp/IpTls
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills:
Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills:
BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills:
Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
Insurance
Designs and operates reliable hybrid application platforms, leading CI/CD, infrastructure automation, cloud resource management, monitoring, security scanning, container orchestration, and database self-service tooling. The role owns the enterprise CI/CD technology stack and hosting strategy across on-premises, hybrid, and cloud environments. It requires extensive collaboration with engineering, infrastructure, security, application, QA, and governance teams, along with 24/7 mission-critical support and technical leadership.
Top Skills:
AnsibleApache CamelAWSCi/CdCloudFormationConfluenceDatabase SystemsDockerGitGithub ActionsGitlab CiGradleHelmInfrastructure As CodeJenkinsJIRAKubernetesLinuxMavenMonitoring And AlertingNetworkingPythonSonatypeTerraformVulnerability Scanning
Artificial Intelligence
Build and maintain reliable infrastructure and software platforms across engineering teams. Configure AWS accounts and networking, Kubernetes clusters, observability and monitoring tools, and CI/CD systems. Write and improve code, support production workloads, participate in on-call operations, estimate tasks, collaborate cross-functionally, and drive assigned reliability initiatives to completion.
Top Skills:
ArgocdAWSCi/CdCoralogixDockerGithub ActionsJavaScriptKubernetesPythonSentryTerraformUnix Shell
Cloud • Security • Software • Cybersecurity
The Site Reliability Engineer II ensures the reliability, availability, performance, and security of critical cloud systems and services. Responsibilities include developing automation for provisioning and configuration management, maintaining monitoring and alerting, optimizing infrastructure performance, supporting high availability, and enabling continuous integration and delivery. The role collaborates with security teams and drives operational improvements across cloud and network infrastructure.
Top Skills:
AnsibleAWSAzureChefContinuous DeliveryContinuous IntegrationDnsElk StackGCPGoGrafanaHTTPKubernetesLinuxPrometheusPuppetPythonShellTcp/IpUnix
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead the design, automation, and scaling of global compute infrastructure across data centers, cloud, and on-prem. Operate GitOps with Rancher Fleet/Flux/Helm, build self-healing tooling, own cluster autoscaling and capacity strategy, define SLOs using Datadog, and participate in on-call rotation while mentoring peers.
Top Skills:
AWSContainerdDatadogDockerFluxGCPGitopsGoHelmHpaInfrastructure As Code (Iac)KarpenterKedaKubernetesLinuxNutanixPythonRancher FleetVsphere
Cloud • Information Technology • Security • Virtual Reality • Cybersecurity
Support reliability, security, scalability, and performance of the TAK CI/CD pipeline and critical infrastructure. Monitor system health, respond to incidents, patch systems, optimize performance and costs, improve DevOps practices, automate operations, and maintain technical documentation. Requires systems and network administration experience, Linux and cloud expertise, containerization, CI/CD knowledge, Zero Trust implementation experience, Security+ certification, and an active T3 investigation.
Top Skills:
AWSCi/Cd PipelinesContainerizationLinuxZero Trust
Cloud • Security • Software • Cybersecurity
Build and maintain reliable, scalable cloud compute platforms across distributed services. Troubleshoot Linux, networking, and production issues; develop automation and AI-assisted tooling; improve monitoring, alerting, SLIs, and SLOs; conduct incident response and root cause analysis; and partner with engineering teams on system design, deployment safety, and operational readiness.
Top Skills:
AnsibleDnsDockerElkGoGrafanaKubernetesLinuxLokiNomadOpensearchPodmanPrometheusPythonSaltTcp/IpTerraform
Artificial Intelligence • Big Data • Information Technology • Other • Software • Database • Biotech
Provides frontline support for high-availability cloud systems, monitors infrastructure, manages incidents, coordinates outage bridges, documents technical issues, and develops automation and AI agentic workflows. The role partners with development teams to improve detection and resolution, supports AWS technologies, databases, networks, CI/CD tools, and observability platforms. This is a Monday–Thursday overnight shift within a 24/7 command center and includes some holiday coverage.
Top Skills:
Ai Agentic WorkflowsAWSAws Application SignalsBashCassandraChefCi/CdCouchdbFirewallsHarnessIntrusion Detection SystemsJavaLan/WanLoad BalancersMssqlMySQLNew RelicPowershellPrompt EngineeringProxy ServersPuppetPythonTcp/IpVirtualization
Automotive
Design and implement scalable cloud infrastructure, monitor performance, automate processes, ensure security and compliance, and lead a DevOps team.
Top Skills:
AWSBashCi/CdDockerElk StackGCPGrafanaKubernetesPrometheusPythonTerraform
Other
Design and operate cloud platforms supporting backend telecom services. Automate deployments, scaling, recovery, and infrastructure provisioning; monitor production systems; maintain observability, alerting, and dashboards; support incident response and on-call operations; manage CI/CD pipelines; and enable engineering, telecom, and data teams through reliable tools and infrastructure.
Top Skills:
AnsibleAWSAzureBashCassandraCircleCICloudFormationDatadogDnsDockerElasticsearchElk StackGitlab CiGoGCPGrafanaHttp/HttpsIamJaegerJenkinsKafkaKubernetesKvmLinuxNoSQLOpentelemetryPerlPrometheusPythonRubySaltstackSplunkSQLTcp/IpTerraformUnixVMware
Artificial Intelligence • HR Tech • Professional Services • Software
Evaluate AI-generated documents, spreadsheets, and presentations for accuracy, rigor, domain quality, and presentation quality in incident management, reliability, and SRE. Apply specialized rubrics, identify factual and aesthetic errors, and provide structured written feedback. This is a fully remote, flexible independent-contractor engagement requiring at least five years of relevant professional experience, English fluency, and proficiency with Microsoft Office and Google Workspace.
Top Skills:
Google SlidesGoogle WorkspaceMS OfficePowerPoint
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own the reliability, performance, resilience, observability, and security of AWS and Kubernetes infrastructure supporting products and AI/ML workloads. Define SLOs, lead incident response and root-cause analysis, build Terraform automation, optimize cloud costs, reduce operational toil, and establish deployment standards that help engineers ship reliably. Participate in on-call rotations and maintain HIPAA-compliant infrastructure.
Top Skills:
AWSClaudeDatadogGitlabGoHipaaIstioKubernetesNatsPostgresPythonSoc 2TerraformTypescript
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Boston Companies Hiring Senior Site Reliability Engineers
See AllPopular Job Searches
All Filters
Total selected ()
No Results
No Results













.png)



.png)











