Maximum of 25 job preferences reached.
Top Senior Site Reliability Engineer Jobs in Boston, MA
Cloud • Security • Software • Cybersecurity
Design, build, and operate scalable infrastructure and CI/CD/IaC systems. Implement observability (monitoring, logging, alerting), automate reliability improvements, mentor engineers, collaborate on incident response, and participate in on-call rotations to maintain Akamai Cloud services.
Top Skills:
AlertingAnsibleBashChefCi/CdGithub ActionsGitlab Ci/CdGoInfrastructure As CodeJenkinsLoggingMonitoringPuppetPythonSaltstackTelemetryTerraform
Artificial Intelligence
Build and maintain reliable infrastructure and software platforms across engineering teams. Configure AWS accounts and networking, Kubernetes clusters, observability and monitoring tools, and CI/CD systems. Write and improve code, support production workloads, participate in on-call operations, estimate tasks, collaborate cross-functionally, and drive assigned reliability initiatives to completion.
Top Skills:
ArgocdAWSCi/CdCoralogixDockerGithub ActionsJavaScriptKubernetesPythonSentryTerraformUnix Shell
Cloud • Security • Software • Cybersecurity
The Site Reliability Engineer II ensures the reliability, availability, performance, and security of critical cloud systems and services. Responsibilities include developing automation for provisioning and configuration management, maintaining monitoring and alerting, optimizing infrastructure performance, supporting high availability, and enabling continuous integration and delivery. The role collaborates with security teams and drives operational improvements across cloud and network infrastructure.
Top Skills:
AnsibleAWSAzureChefContinuous DeliveryContinuous IntegrationDnsElk StackGCPGoGrafanaHTTPKubernetesLinuxPrometheusPuppetPythonShellTcp/IpUnix
Cloud • Information Technology • Security • Virtual Reality • Cybersecurity
Support reliability, security, scalability, and performance of the TAK CI/CD pipeline and critical infrastructure. Monitor system health, respond to incidents, patch systems, optimize performance and costs, improve DevOps practices, automate operations, and maintain technical documentation. Requires systems and network administration experience, Linux and cloud expertise, containerization, CI/CD knowledge, Zero Trust implementation experience, Security+ certification, and an active T3 investigation.
Top Skills:
AWSCi/Cd PipelinesContainerizationLinuxZero Trust
Cloud • Security • Software • Cybersecurity
Build and maintain reliable, scalable cloud compute platforms across distributed services. Troubleshoot Linux, networking, and production issues; develop automation and AI-assisted tooling; improve monitoring, alerting, SLIs, and SLOs; conduct incident response and root cause analysis; and partner with engineering teams on system design, deployment safety, and operational readiness.
Top Skills:
AnsibleDnsDockerElkGoGrafanaKubernetesLinuxLokiNomadOpensearchPodmanPrometheusPythonSaltTcp/IpTerraform
Cloud • Security • Software • Cybersecurity
Lead site reliability engineering for Akamai’s compute infrastructure and services. Develop automation, define reliability requirements and standards, establish SLOs and KPIs, troubleshoot complex distributed-system and hardware issues, manage incidents and postmortems, and participate in on-call rotations. Coordinate engineering teams, support strategic initiatives, mentor engineers, and guide restoration of service-impacting issues.
Top Skills:
Cloud InfrastructureDistributed SystemsInfrastructure AutomationLinux
Automotive
Design and implement scalable cloud infrastructure, monitor performance, automate processes, ensure security and compliance, and lead a DevOps team.
Top Skills:
AWSBashCi/CdDockerElk StackGCPGrafanaKubernetesPrometheusPythonTerraform
Other
Design and operate cloud platforms supporting backend telecom services. Automate deployments, scaling, recovery, and infrastructure provisioning; monitor production systems; maintain observability, alerting, and dashboards; support incident response and on-call operations; manage CI/CD pipelines; and enable engineering, telecom, and data teams through reliable tools and infrastructure.
Top Skills:
AnsibleAWSAzureBashCassandraCircleCICloudFormationDatadogDnsDockerElasticsearchElk StackGitlab CiGoGCPGrafanaHttp/HttpsIamJaegerJenkinsKafkaKubernetesKvmLinuxNoSQLOpentelemetryPerlPrometheusPythonRubySaltstackSplunkSQLTcp/IpTerraformUnixVMware
Artificial Intelligence • HR Tech • Professional Services • Software
Evaluate AI-generated documents, spreadsheets, and presentations for accuracy, rigor, domain quality, and presentation quality in incident management, reliability, and SRE. Apply specialized rubrics, identify factual and aesthetic errors, and provide structured written feedback. This is a fully remote, flexible independent-contractor engagement requiring at least five years of relevant professional experience, English fluency, and proficiency with Microsoft Office and Google Workspace.
Top Skills:
Google SlidesGoogle WorkspaceMS OfficePowerPoint
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own the reliability, performance, resilience, observability, and security of AWS and Kubernetes infrastructure supporting products and AI/ML workloads. Define SLOs, lead incident response and root-cause analysis, build Terraform automation, optimize cloud costs, reduce operational toil, and establish deployment standards that help engineers ship reliably. Participate in on-call rotations and maintain HIPAA-compliant infrastructure.
Top Skills:
AWSClaudeDatadogGitlabGoHipaaIstioKubernetesNatsPostgresPythonSoc 2TerraformTypescript
Information Technology • Internet of Things • Software • Virtual Reality
Lead reliability, scalability, and operational excellence for large-scale Java distributed systems. Drive architecture improvements, incident management, observability, SLO adoption, automation, infrastructure strategy, and performance optimization. Influence cross-team technical decisions, mentor engineers, establish reliability standards, and guide incident response and durable remediation. The role requires hybrid work in Boston and participation in an on-call rotation.
Top Skills:
AWSCi/CdInfrastructure As CodeJavaMongoDBRabbitMQZookeeper
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Lead architecture and build a Kubernetes-based, GitOps-driven platform and self-service datastore offerings. Drive IaC, observability, and platform automation; partner with application teams to diagnose and optimize datastore and messaging performance at scale.
Top Skills:
AWSAzureCassandraDatadogGCPGitopsGrafanaKafkaKubernetesLlm/Agentic Ai ToolingMySQLNew RelicPostgresPulumiTerraform
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Reposted One Month AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills:
AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
Owns reliability, observability, and incident response for a GPUaaS platform. Defines SLOs, builds monitoring and alerting systems, leads major incidents and post-incident reviews, automates operational processes, maintains runbooks, manages on-call operations, coordinates with engineering teams, drives chaos testing, reports SLA performance, and mentors junior engineers.
Top Skills:
DatadogGoGpuaasGrafanaGremlinHpcKubernetesLitmusOpentelemetryPrometheusPython
Software
Lead SRE responsible for incident escalation, platform architecture reviews, availability and SLO management, IaC (Terraform), Kubernetes production operations, hybrid-cloud administration, observability (Datadog), alerting (PagerDuty), automation with scripting, and ensuring compliance (PCI-DSS, SOC1/2, ISO27001). Mentor engineers and drive platform-level remediation and DevOps improvements.
Top Skills:
Active DirectoryAWSBashCi/CdCloudflareDatadogDnsFirewallsKubernetesAzurePagerdutyPowershellRoutingTerraformVpn
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills:
BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Edtech
Lead infrastructure modernization and platform reliability across multiple cloud providers. Design infrastructure as code, operate Kubernetes and Linux environments, improve CI/CD and deployment tooling, establish SLI/SLO practices, strengthen observability, lead incident response, manage cloud costs, and partner on security and compliance. Provide technical leadership through architecture guidance, mentorship, engineering standards, and roadmap development while participating in on-call support.
Top Skills:
AWSCi/CdGCPJenkinsKubernetesLinuxPythonRubyRuby On RailsSoc 2SpinnakerTerraform
Cloud • Information Technology • Business Intelligence • Consulting
Design, build, and operate cloud infrastructure and SRE capabilities for an enterprise AI platform. Responsibilities include infrastructure-as-code, landing zones, networking, Kubernetes, CI/CD, observability, incident response, SLOs, production readiness, automation, cost optimization, and support for hybrid, edge, air-gapped, and customer-controlled environments. The role is remote, client-facing, and requires strong collaboration and reliability ownership.
Top Skills:
AlertingAzureAzure ArcAzure DevopsBicepCi/CdDashboardsDockerGithub ActionsGpu WorkloadsInfrastructure As CodeKubernetesLogsMetricsObservabilitySlosTerraformTraces
Enterprise Web • Hardware • Internet of Things • Software
Lead observability and reliability efforts: mentor teams on SLIs/SLOs, maintain triage/remediation workflows, perform incident response, debug production systems, and design core infrastructure and tooling for engineering teams.
Top Skills:
AlloyClaude SkillsGemini GemsGoGrafanaKubernetesLokiMimirMongoDBOpentelemetryPostgresPrometheusPromqlTempoTypescript
Artificial Intelligence • Cloud • Information Technology • Software • Big Data Analytics
Own reliability for Kong’s Volcano internal developer platform by defining SLOs, incident practices, and observability. Design multi-region Kubernetes infrastructure, GitOps deployment automation, preview environments, managed PostgreSQL, Redis, and object storage. Lead reliability and compliance initiatives with engineering, security, and OCTO leadership while evaluating emerging edge, serverless, vector database, and AI infrastructure technologies.
Top Skills:
ArgocdAutoscalingCi/CdCniDatadogGitopsGrafanaHelmIngressKubernetesObject StoragePostgresPrometheusRedisService MeshTerraformTerragrunt
Artificial Intelligence • Healthtech • Software • Telehealth
Designs, deploys, and maintains resilient AWS and Kubernetes infrastructure. Builds automation, GitHub Actions components, internal AI-assisted operational tools, and observability systems. Leads incident response, postmortems, and SLO/SLI management while ensuring HIPAA compliance and high availability. Collaborates across teams on architecture reviews, risk reduction, clinical safety, and reliability best practices.
Top Skills:
Ai-Assisted OperationsAmazon Ec2Amazon EksAmazon RdsAmazon S3AWSBashDatadogGithub ActionsGoHelmKubernetesPythonTerraform
Cloud • Security • Software • Cybersecurity
As a Site Reliability Engineer II, you'll automate tasks, monitor AI workloads, enhance dashboards, support CI/CD processes, and collaborate with engineering teams on complex issues while participating in on-call rotations.
Top Skills:
GoGrafanaKubernetesLinuxPrometheusPythonSaltstackTerraform
Cloud • Information Technology • Biotech
The Site Reliability Engineer will build and deploy Linux servers, research technologies, monitor system performance, and resolve technical incidents.
Top Skills:
Infrastructure-As-CodeLinuxNetworkingVirtualization
Artificial Intelligence • Healthtech • Software
Build and maintain scalable infrastructure platforms supporting domestic and international workloads. Develop CI/CD, declarative lifecycle management, Kubernetes clusters, monitoring, and automation systems. Troubleshoot infrastructure issues, respond to alerts, minimize downtime, streamline delivery pipelines and database changes, and help guide SRE team direction. Collaborate with engineers, data scientists, and technology professionals while promoting a high-performance, cross-functional culture.
Top Skills:
AWSAzureContainerdContinuous DeploymentContinuous IntegrationDnsDockerFirewallsGCPGoGrpcHelmKubernetesLinuxLoad BalancingPrometheusPythonRoutingShell ScriptingTcp/IpUdp
Software
Operate and maintain highly available, secure, containerized SaaS applications across AWS and Azure. Responsibilities include rotating 24x7 on-call coverage, observability, incident and security response, disaster recovery, Terraform infrastructure-as-code, CI/CD automation, cloud integration, performance optimization, and developing AI agents to automate SRE and DevSecOps workflows.
Top Skills:
Amazon Web ServicesCheckovClaude CodeDatadogDockerDynatraceGithub ActionsGitlab CiGoJavaScriptKubernetesAzureNew RelicPrisma CloudPythonTerraformWiz
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Boston Companies Hiring Senior Site Reliability Engineers
See AllPopular Job Searches
All Filters
Total selected ()
No Results
No Results






.png)
















.jpg)








