Top SRE Engineer Jobs in Boston, MA

16 Days AgoSaved
In-Office or Remote
Boston, MA
Mid level
Mid level
Artificial Intelligence • Fintech • Machine Learning • Software • App development • Conversational AI • Generative AI
Own and improve production infrastructure reliability, deployments, Infrastructure-as-Code, Kubernetes environments, automation, CI/CD, monitoring, alerting, and observability. Investigate incidents, optimize system performance, maintain documentation and runbooks, and support DNS, WAF, CDN, and caching infrastructure. The role requires strong Linux administration, Bash scripting, networking, Git, and containerization skills, with independent ownership and collaboration across development and operations teams.
Top Skills: AkamaiAmqpAnsibleAWSBashCdnCloudflareDnsDockerGCPGitGitlab CiGrafanaHttp/HttpsKubernetesLinuxPodmanPrometheusPythonRabbitMQTerraformVictoriametricsWafZabbix
Reposted 16 Days AgoSaved
In-Office or Remote
Boston, MA
200K-200K Annually
Mid level
200K-200K Annually
Mid level
Cloud • Software
The Site Reliability Engineer will ensure reliable cloud operations by applying Python for infrastructure automation, managing OpenStack and Kubernetes, and practicing devsecops in a fast-paced environment.
Top Skills: KubernetesLinuxOpenstackPython
Reposted 16 Days AgoSaved
In-Office or Remote
Boston, MA
165K-215K Annually
Senior level
165K-215K Annually
Senior level
Software • Cybersecurity
This role involves managing Kubernetes clusters, cloud infrastructure, and CI/CD pipelines. The engineer will enhance system reliability and efficiency while troubleshooting production issues.
Top Skills: AlertmanagerAWSAzureBashCi/CdDockerElastic StackElasticsearchGCPGoGrafanaHelmKafkaKubernetesLokiMongoDBOciPrometheusPythonRedisSparkTerraform
Reposted 16 Days AgoSaved
Remote
Boston, MA
Senior level
Senior level
Software • Web3
Lead reliability practices across teams: embed early in projects, define SLIs/SLOs, build multi-cloud paved roads with Terraform, run on-call, drive org-wide incident maturity and tooling.
Top Skills: AWSAzureGCPRuby On RailsTerraformTypescriptWebcontainers
27 Days AgoSaved
In-Office
Boston, MA
Senior level
Senior level
Information Technology • Productivity • Software • Manufacturing
Senior individual contributor SRE responsible for reliability, observability, automation, and incident command across multi-account AWS SaaS. Design Terraform modules, CI/CD pipelines, autoscaling/self-healing, OpenTelemetry-based monitoring, and security controls; mentor engineers and integrate validated AI tooling.
Top Skills: Aws Ec2CloudwatchEcs FargateEksGithub ActionsIamLambdaOpenobserveOpentelemetryPagerdutyPythonRds PostgresqlS3TerraformVpc
Reposted 17 Days AgoSaved
Remote
Boston, MA
120K-165K Annually
Senior level
120K-165K Annually
Senior level
Fitness • Healthtech • Software
Own reliability and security of CI/CD and production services: define SLI/SLOs, lead incident response and postmortems, build observability (Datadog), operate Kubernetes and IaC (Terraform), harden pipelines with SAST/DAST/SCA and policy-as-code, and coach teams on reliability and operational best practices.
Top Skills: AWSCi/CdConftestDastDatadogDockerGithub ActionsGoInfrastructure As CodeKubernetesKyvernoOpa/RegoPythonSastScaTerraformTypescript
18 Days AgoSaved
In-Office or Remote
Boston, MA
142K-268K Annually
Expert/Leader
142K-268K Annually
Expert/Leader
Automotive
Leads SRE engineering leaders and engineers while defining enterprise observability, reliability, and platform strategy across GCP, on-premise, manufacturing, distribution, and campus environments. Oversees vendor-agnostic tooling, OpenTelemetry integrations, CI/CD observability, SRE maturity models, and Agentic AI initiatives. Drives adoption of SRE practices, develops technical roadmaps, partners with senior leadership and operational teams, and maintains hands-on architectural and technical credibility.
Top Skills: Agentic AiAWSAzureCi/CdDatadogDynatraceGCPNew RelicOpentelemetryOtel Genai Semantic ConventionsSource Control PlatformsSplunkTerraform
Reposted 27 Days AgoSaved
In-Office or Remote
Boston, MA
146K-264K Annually
Senior level
146K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Ensure reliability, scalability, and usability of network infrastructure for Akamai Connected Cloud. Define requirements and SLOs, build automation and CI/CD pipelines, collaborate with dev/QA to improve code and stability, troubleshoot complex network issues (on-call), and mentor teammates while driving architectural standards.
Top Skills: AnsibleArgocdBashBirdChefFrrGithub ActionsGoGobgpJenkinsLinux NetworkingPuppetPythonSalt Stack
Reposted 18 Days AgoSaved
Remote
Boston, MA
130K-160K Annually
Senior level
130K-160K Annually
Senior level
Other
Design, build, and maintain highly available cloud-native systems. Improve reliability through automation, CI/CD, Kubernetes, observability, and incident management. Collaborate with developers, security, and product teams to define SLOs, implement self-healing, debug production issues, and ensure secure deployments.
Top Skills: AWSAzure Cloud ServicesDatadogGCPGithub ActionsGitlab CiGoInfrastructure As CodeKubernetesOpsgeniePagerdutyPythonRubySite Reliability Engineering Foundation
Reposted 18 Days AgoSaved
In-Office or Remote
Boston, MA
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Software
Design and operate large-scale GPU infrastructure for distributed AI training, ensuring reliability, performance, and efficient customer partnerships.
Top Skills: AnsibleCudaDeepspeedFsdpGpuHelmInfinibandKubernetesLinuxMegatronNcclNvidia A100Nvidia B200Nvidia H100NvlinkPyTorchRoceTerraform
Reposted 18 Days AgoSaved
Remote
Boston, MA
180K-224K Annually
Senior level
180K-224K Annually
Senior level
Artificial Intelligence • Information Technology • Consulting
Build and operate Nebius's network infrastructure: define SLIs/SLOs, improve site and inter-site reliability, lead incident response and postmortems, develop observability and alerting, automate change workflows, and collaborate with network and platform teams to embed operability.
Top Skills: Ci/CdContainer PlatformsGoInfrastructure As CodeLinuxPython
Reposted 20 Days AgoSaved
In-Office or Remote
Boston, MA
Senior level
Senior level
Software
The role involves managing compute infrastructure for decentralized applications, requiring critical thinking, documentation skills, and experience in Kubernetes and blockchain management.
Top Skills: BlockchainGitopsInfrastructure-As-CodeKubernetesProgramming Languages
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted 20 Days AgoSaved
Remote
Boston, MA
150K-195K Annually
Senior level
150K-195K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Database
As a Site Reliability Engineer, you will design, implement, and maintain scalable infrastructure, ensure system reliability, automate processes, and collaborate with engineering teams.
Top Skills: DockerElk StackGoGrafanaJavaKubernetesNode.jsPrometheusPulumiPythonRubyTerraform
Reposted 20 Days AgoSaved
Remote
Boston, MA
117K-181K Annually
Senior level
117K-181K Annually
Senior level
Other • Social Impact
As a Senior Site Reliability Engineer, you will design, develop, and maintain reliable infrastructure for Wikimedia's API services, ensuring performance and availability while driving reliability engineering practices and improving developer experience.
Top Skills: AnsibleArgocdAWSAzureGCPGitlabGoKubernetesOpentelemetryPrometheusPythonTerraform
21 Days AgoSaved
Remote
Boston, MA
Entry level
Entry level
Angel or VC Firm • Blockchain • Fintech • Cryptocurrency
Apply to join Galaxy Ventures' invite-only Talent Network for DevOps, SRE, QA, and Security professionals. Upon acceptance, your profile may be discreetly shared with portfolio companies for relevant roles, and you'll receive invitations to exclusive networking events. Participation is confidential and non-binding.
Reposted One Month AgoSaved
In-Office or Remote
Boston, MA
146K-264K Annually
Senior level
146K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Lead and mentor SRE teams; partner with engineering, operations and product; apply statistical analysis and networking expertise to diagnose performance and reliability issues; define and implement data feeds; influence technical decisions and investments; build tooling to automate analytical workflows and increase platform reliability.
Top Skills: CCloudDistributed SystemsDnsEdgeHTTPJavaPerlPythonRSQLTcpTls
Reposted 21 Days AgoSaved
Remote
Boston, MA
160K-208K Annually
Senior level
160K-208K Annually
Senior level
Healthtech • Software
Design, automate, and maintain scalable infrastructure and SRE tooling. Manage Kubernetes clusters, CI/CD, monitoring, and incident response. Improve processes, reduce toil via automation, and collaborate with engineering and data teams to support domestic and international workloads.
Top Skills: AWSAzureContainerdDnsDockerFirewallsGCPGoGrpcHelmKubernetesLinuxLoad BalancingPrometheusPythonRoutingShell ScriptingTcp/IpUdp
Reposted 22 Days AgoSaved
Remote
Boston, MA
180K-210K Annually
Senior level
180K-210K Annually
Senior level
Artificial Intelligence • Insurance • Software • Automation
The Staff Site Reliability Engineer will build and scale infrastructure for Assured's platform, automate delivery, enhance observability, and lead mentoring initiatives.
Top Skills: AWSKubernetesPostgresTerraform
Reposted 22 Days AgoSaved
Remote
Boston, MA
Senior level
Senior level
Healthtech • Social Impact • Software
Own the operational lifecycle of cloud-native data infrastructure: design and automate reliable deployments, observability, incident response, SLIs/SLOs, autoscaling and IaC, and improve platform efficiency and data freshness across GKE and Cloud Run.
Top Skills: BashBigQueryCloud BuildCloud MonitoringCloud RunDatadogDockerGCPGithub ActionsGkeGoGrafanaJIRAKubernetesPrometheusPulumiPythonSentrySlackSnykSonarqubeTerraform
Reposted 22 Days AgoSaved
Remote
Boston, MA
205K-270K Annually
Senior level
205K-270K Annually
Senior level
Artificial Intelligence • Other • Sales • Software
The role involves designing and advancing infrastructure for the engineering team, ensuring the reliability of Kubernetes clusters, automating operations, and building machine learning infrastructure.
Top Skills: ArgoAWSAzureCloudFormationFluxGithub ActionsGoGCPKubernetesPostgresPythonTerraform
Reposted 22 Days AgoSaved
Remote
Boston, MA
Senior level
Senior level
Software • Consulting
Lead 24x7 application support for external web applications: manage incidents, perform RCA, implement preventative fixes, build monitoring/alerts, expand Splunk functionality, create dashboards, collaborate with development and platform teams, and participate in on-call rotation.
Top Skills: ApmAppdynamicsAWSDatadogGrafanaKubernetesLinuxMulesoftOpenshiftOpentelemetryPostmanPythonRumSeleniumServicenowShell ScriptingSplunk (Spl)Splunk CloudSplunk Observability CloudSplunk Synthetics
Reposted 22 Days AgoSaved
Remote
Boston, MA
160K-180K Annually
Senior level
160K-180K Annually
Senior level
Software
Own and improve platform performance, reliability, and deployment automation. Manage cloud infrastructure, implement IaC, monitor systems with observability tools, provide operational support for distributed applications, and integrate production learnings into development workflows.
Top Skills: Aiops ToolingAws Elastic ContainersAws RdsAws S3Claude CodeClaude CoworkDatadogHarness EngineeringInfrastructure As CodeKubernetesLlmsPrompt EngineeringRigorSplunk
23 Days AgoSaved
Remote
Boston, MA
Senior level
Senior level
Information Technology • Professional Services
Provide senior SRE expertise to improve reliability, scalability, performance, and resilience of a cloud-hosted geospatial platform. Design monitoring/observability, automate deployments, support incident response, optimize capacity and performance, and collaborate across DevSecOps, Kubernetes, database, and support teams in a mission-focused DoD environment.
Top Skills: Alerting ToolsArcgis EnterpriseAw S Cloud OneAWSCi/CdContainerized SystemsEsriInfrastructure-As-CodeKubernetesLinuxLogging ToolsMonitoring ToolsRmfScriptingStig
One Month AgoSaved
In-Office
Boston, MA
181K-237K Annually
Senior level
181K-237K Annually
Senior level
Healthtech • Insurance
Design, build, and operate resilient cloud infrastructure and platform services. Lead cross-team projects, mentor engineers, define SLOs, automate CI/CD and IaC, improve reliability, and own medium-to-large infrastructure features.
Top Skills: ArgocdAWSGCPGithub ActionsGrafanaIamIstioKubernetesPrometheusTerraformVpc Peering
Reposted 23 Days AgoSaved
Remote
Boston, MA
175K-195K Annually
Senior level
175K-195K Annually
Senior level
Legal Tech • Software
Design and improve observability (monitoring, logging, tracing, SLIs/SLOs), build automation and CI/CD, lead incident response and reliability improvements, mentor SREs, run on-call, and apply AI/ML to operational signals to forecast and reduce risks.
Top Skills: AWSBashCi/CdDistributed TracingGoInfrastructure As CodeKubernetesLoggingMonitoringPythonSlisSlos
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account