Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in Boston, MA
Software
Owns reliability, observability, performance, and security for a multi-region SaaS platform. Responsibilities include managing Datadog, implementing APM and tracing, defining SLOs, developing automation, expanding infrastructure as code and CI/CD, automating operational workflows, maintaining security controls, participating in incident response, documenting procedures, and mentoring engineers.
Top Skills:
ApmAzure DevopsAzure Kubernetes ServiceAzure SqlBashBicepCi/CdCosmos DbDatadogDistributed TracingHelmInfrastructure As CodeKey VaultKubernetesKustomizeManaged IdentitiesAzureMicrosoft Entra IdPowershellPythonRedisService BusTerraform
Fintech • Software
Senior SRE responsible for ensuring reliability, scalability, and performance of production systems. Investigates and resolves incidents, works with R&D on defects, manages deployments and change validation, builds monitoring and diagnostic tooling to improve MTTA/MTTD/MTTR, configures observability platforms, participates in on-call rotation, and performs risk assessments and production readiness validation.
Top Skills:
Ai ToolsAkamaiAmqAnsibleAWSAzureBashCdnCloudflareCloudwatchDatadogDatastreamDnsDockerDynatraceEfsEksHTTPHttpsInterconnectJavaJmsKubernetesMongoDBOraclePostgresPrometheusPythonRabbitMQRedshiftS3SplunkTcpTerraformUdpUnix/LinuxWafZabbix
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills:
BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Edtech
Lead infrastructure modernization and platform reliability across multiple cloud providers. Design infrastructure as code, operate Kubernetes and Linux environments, improve CI/CD and deployment tooling, establish SLI/SLO practices, strengthen observability, lead incident response, manage cloud costs, and partner on security and compliance. Provide technical leadership through architecture guidance, mentorship, engineering standards, and roadmap development while participating in on-call support.
Top Skills:
AWSCi/CdGCPJenkinsKubernetesLinuxPythonRubyRuby On RailsSoc 2SpinnakerTerraform
Cloud • Information Technology • Business Intelligence • Consulting
Design, build, and operate cloud infrastructure and SRE capabilities for an enterprise AI platform. Responsibilities include infrastructure-as-code, landing zones, networking, Kubernetes, CI/CD, observability, incident response, SLOs, production readiness, automation, cost optimization, and support for hybrid, edge, air-gapped, and customer-controlled environments. The role is remote, client-facing, and requires strong collaboration and reliability ownership.
Top Skills:
AlertingAzureAzure ArcAzure DevopsBicepCi/CdDashboardsDockerGithub ActionsGpu WorkloadsInfrastructure As CodeKubernetesLogsMetricsObservabilitySlosTerraformTraces
Healthtech • Insurance
Lead cloud, DevOps, and SRE architecture efforts to scale CareSource's digital platform. Partner with Cloud, DevOps, Security, and SRE teams to design infrastructure, build Terraform templates, enhance CI/CD pipelines, implement monitoring/alerting, and improve reliability, scalability, and incident response for enterprise-scale digital products.
Top Skills:
Azure CloudDockerDynatraceGithub ActionsKubernetesSplunkTerraform Enterprise
Reposted One Month AgoSaved
Easy Apply
Easy Apply
Cloud • Information Technology • Security • Software • Cybersecurity
This internship role focuses on SRE skills, requiring collaboration and problem-solving in dynamic environments for Zscaler's Zero Trust Exchange team.
Top Skills:
AnsibleAws EcsKubernetesLinuxPythonTerraform
Artificial Intelligence • Healthtech • Software • Telehealth
Designs, deploys, and maintains resilient AWS and Kubernetes infrastructure. Builds automation, GitHub Actions components, internal AI-assisted operational tools, and observability systems. Leads incident response, postmortems, and SLO/SLI management while ensuring HIPAA compliance and high availability. Collaborates across teams on architecture reviews, risk reduction, clinical safety, and reliability best practices.
Top Skills:
Ai-Assisted OperationsAmazon Ec2Amazon EksAmazon RdsAmazon S3AWSBashDatadogGithub ActionsGoHelmKubernetesPythonTerraform
Cloud • Security • Software • Cybersecurity
As a Site Reliability Engineer II, you'll automate tasks, monitor AI workloads, enhance dashboards, support CI/CD processes, and collaborate with engineering teams on complex issues while participating in on-call rotations.
Top Skills:
GoGrafanaKubernetesLinuxPrometheusPythonSaltstackTerraform
Cloud • Information Technology • Biotech
The Site Reliability Engineer will build and deploy Linux servers, research technologies, monitor system performance, and resolve technical incidents.
Top Skills:
Infrastructure-As-CodeLinuxNetworkingVirtualization
Hardware • Quantum Computing
Lead integration, maintenance, and automation of heterogeneous hardware and software control systems for quantum computers. Manage networked lab infrastructure, CI/CD pipelines, observability, and provisioning. Support incident response, testing, and orchestration, collaborating with software, hardware, and test teams to ensure reliability and operational readiness of development and production environments.
Top Skills:
AnsibleBashCi/CdDebianDhcpDnsDockerElkGitGitlab CiGoGrafanaHardware-In-The-Loop (Hil)JenkinsKubernetesLanLogging SystemsPrometheusPythonRack-Mount ServersRed HatRoutersSwitchesTcp/IpTerraformUbuntuVlanWanWindows
Cloud • Security • Software • Cybersecurity
Design, build, and operate scalable infrastructure and CI/CD/IaC systems. Implement observability (monitoring, logging, alerting), automate reliability improvements, mentor engineers, collaborate on incident response, and participate in on-call rotations to maintain Akamai Cloud services.
Top Skills:
AlertingAnsibleBashChefCi/CdGithub ActionsGitlab Ci/CdGoInfrastructure As CodeJenkinsLoggingMonitoringPuppetPythonSaltstackTelemetryTerraform
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Cloud • Security • Software • Cybersecurity
Design, develop, test, and operate scalable infrastructure and services for Akamai Cloud. Implement and manage Infrastructure-as-Code (Terraform and similar tools), CI/CD, and observability. Automate reliability improvements, mentor engineers, collaborate on incident response and root-cause remediation, and participate in on-call rotations.
Top Skills:
Alerting)AnsibleChefCi/CdInfrastructure As CodeLinuxLoggingObservability (MonitoringPuppetSaltstackTerraform
Gaming
Own and operate large-scale infrastructure for sports betting and media platforms across cloud and production environments. Lead infrastructure migrations, build Kubernetes platform tooling and CI/CD automation, improve observability and alerting, support development teams, and participate in incident response. The role requires strong distributed-systems expertise, production troubleshooting, cross-team project leadership, technical communication, and mentoring.
Top Skills:
ArgocdAWSBashCephCiliumDatadogGCPGithub ActionsGoHelmIstioKubernetesLinuxPgbouncerPostgresPythonTalos OsTerraform
Digital Media
Build, maintain, and operate Ookla’s globally distributed infrastructure platform at massive scale. Responsibilities include managing cloud instances, containers, serverless applications, databases, streaming systems, and big-data tooling; supporting 24/7 production operations and on-call rotations; implementing security programs; improving deployment pipelines, monitoring, observability, and reliability; and guiding software and data engineering teams on operational best practices and troubleshooting.
Top Skills:
Amazon AuroraAmazon RdsAnsibleSparkAWSChefCloudFormationDockerDynamoDBGitGitGoIds/IpsJavaKafkaKinesisKubernetesLinuxMongoDBMySQLPHPPostgresPythonRubySQLTerraformTypescript
Aerospace • Manufacturing
Build and lead a centralized observability platform for satellite, ground-station, and distributed network systems. Responsibilities include scaling metrics, logging, and tracing infrastructure; defining SLOs, SLIs, and error budgets; enabling application instrumentation; automating deployments with Terraform and ArgoCD; monitoring Kubernetes, GCP, and AWS environments; and developing incident response, alerting, and reliability practices. The role includes on-call responsibilities and requires an active Top Secret/SCI clearance.
Top Skills:
ArgocdAWSC++ElkGitlab CiGoGoogle Cloud PlatformGrafanaHoneycombIstioJaegerJavaKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
Cloud • Security • Software
Design, deploy, and maintain resilient cloud infrastructure for Ping Identity’s mission-critical services. Build and optimize automated CI/CD pipelines, support cloud security and observability, evaluate technologies, participate in planning and on-call rotations, and help improve engineering practices. Collaborate across development and operations teams while sharing expertise and supporting distributed production systems.
Top Skills:
Ci/CdCloud PlatformsDistributed SystemsDockerGitGoIdentity And Access ManagementKubernetesNetworking
Healthtech • Software
Provides technical leadership for site reliability and operational development. Designs automated, repeatable cloud and infrastructure solutions; monitors service objectives and business metrics; improves application resiliency, performance, efficiency, and cost. Builds monitoring, deployment, testing, and vulnerability-response automation while supporting mission-critical production systems. Collaborates with software, security, product, and business teams, conducts technical training and resilience exercises, and promotes sound development, change-management, and operational practices.
Top Skills:
Amazon EcsAnsibleAzureCC++ChefDockerGoJavaKubernetesLinuxPerlPuppetPythonRubyTerraformWindows
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Lead architecture and build a Kubernetes-based, GitOps-driven platform and self-service datastore offerings. Drive IaC, observability, and platform automation; partner with application teams to diagnose and optimize datastore and messaging performance at scale.
Top Skills:
AWSAzureCassandraDatadogGCPGitopsGrafanaKafkaKubernetesLlm/Agentic Ai ToolingMySQLNew RelicPostgresPulumiTerraform
Artificial Intelligence • Software
Architects and owns highly available infrastructure and Kubernetes-based platforms supporting autonomous systems. Builds Golang backend services, platform tooling, observability systems, dashboards, alerts, and log aggregation. Partners with product teams to launch services, performs performance analysis, manages cloud upgrades, and participates in incident response and postmortems. Collaborates on cloud security risk assessments, intrusion detection, threat-feed systems, risk mitigation, and SaaS payment processes. Provides architectural leadership and mentorship across engineering teams.
Top Skills:
ArgocdArgocd Image UpdaterArtifactoryAWSGithub ActionsGoJavaScriptKubernetesPythonRustTerraform
Information Technology • Security • Cybersecurity
Own production operations and highly available cloud infrastructure for regulated government environments. Lead incident response, root-cause analysis, disaster recovery testing, compliance operationalization, audit readiness, vulnerability management, and continuous monitoring. Build automation, secure CI/CD pipelines, infrastructure-as-code, observability, and compliance tooling across Kubernetes, Linux, containers, and cloud platforms. Partner with security, compliance, and engineering teams to improve reliability, deployment safety, and regulatory sustainment.
Top Skills:
Aws GovcloudBashCi/CdDod Il4Dod Il5FedrampGitopsGoGrafanaKubernetesLinuxNist 800-53PrometheusPythonStigTerraformUnixZero Trust
Cloud • Security • Software • Cybersecurity
Lead reliability for a serverless AI inference platform: own observability and SLO/SLI frameworks, build automation and tooling, manage incidents and on-call, define deployment safety (canaries, rollbacks), influence architecture with product teams, and mentor other SREs.
Top Skills:
AutoscalingCi/CdContainer OrchestrationContainerizationGoGpu WorkloadsInfrastructure-As-CodeKubernetesModel ServingPythonResource Scheduling
Aerospace • Manufacturing
As a Site Reliability Engineer, you'll build and manage observability platforms for satellite communications, define SLOs/SLIs, and collaborate on incident response and deployment automation.
Top Skills:
ArgocdAWSElkGCPGoGrafanaIstioJaegerKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
Cybersecurity
Assist with monitoring system performance, uptime, and reliability; support incident response, troubleshooting, and root cause analysis; monitor SRE alerts in Slack and report issues; learn cloud platforms; and collaborate with development and DevOps teams.
Top Skills:
AWSAzureDevOpsDevsecopsGCPGrafanaKubernetesPrometheusSlack
Artificial Intelligence • Fintech • Machine Learning • Natural Language Processing • Business Intelligence
Lead architecture and implementation of reliability platforms and SRE practices for a production SaaS. Build self-service reliability tooling, drive AIOps automation, advance observability (monitoring, tracing, profiling), lead incident response and postmortems, mentor engineers, and embed production readiness across teams to achieve 99.99% uptime.
Top Skills:
AWSAzureContinuous ProfilingDatadogDnsElkGCPGoGrafanaHttp/SKubernetesLoad BalancingOpentelemetryPrometheusPythonTcp/Ip
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Boston, MA Companies Hiring SRE Engineers
See AllPopular Boston, MA Engineering Job Searches
Engineering Jobs in Boston, MA
Software Engineer Jobs in Boston, MA
Android Developer Jobs in Boston, MA
C# Jobs in Boston, MA
C++ Jobs in Boston, MA
DevOps Jobs in Boston, MA
Front End Developer Jobs in Boston, MA
Golang Jobs in Boston, MA
Hardware Engineer Jobs in Boston, MA
iOS Developer Jobs in Boston, MA
Java Developer Jobs in Boston, MA
Javascript Jobs in Boston, MA
Linux Jobs in Boston, MA
Engineering Manager Jobs in Boston, MA
.NET Developer Jobs in Boston, MA
PHP Developer Jobs in Boston, MA
Python Jobs in Boston, MA
QA Jobs in Boston, MA
Ruby Jobs in Boston, MA
Salesforce Developer Jobs in Boston, MA
Scala Jobs in Boston, MA
Application Engineer Jobs in Boston, MA
Automation Engineer Jobs in Boston, MA
AWS Engineer Jobs in Boston, MA
Backend Engineer Jobs in Boston, MA
Cloud Engineer Jobs in Boston, MA
Controls Engineer Jobs in Boston, MA
CTO Jobs in Boston, MA
Design Engineer Jobs in Boston, MA
DevOps Engineer Jobs in Boston, MA
Director of Engineering Jobs in Boston, MA
Electrical Engineering Jobs in Boston, MA
Embedded Software Engineer Jobs in Boston, MA
Full-Stack Engineer Jobs in Boston, MA
Game Engineer Jobs in Boston, MA
Infrastructure Engineer Jobs in Boston, MA
Manufacturing Engineer Jobs in Boston, MA
Mechanical Design Engineer Jobs in Boston, MA
Mechanical Engineering Jobs in Boston, MA
Mechatronics Engineering Jobs in Boston, MA
Network Engineer Jobs in Boston, MA
Platform Engineer Jobs in Boston, MA
Principal Engineer Jobs in Boston, MA
Principal Software Engineer Jobs in Boston, MA
Process Engineer Jobs in Boston, MA
Product Engineer Jobs in Boston, MA
Project Engineer Jobs in Boston, MA
QA Engineer Jobs in Boston, MA
Robotics Engineer Jobs in Boston, MA
Security Engineer Jobs in Boston, MA
Software Engineering Manager Jobs in Boston, MA
Software Test Engineer Jobs in Boston, MA
Solutions Architect Jobs in Boston, MA
Solutions Engineer Jobs in Boston, MA
SRE Engineer Jobs in Boston, MA
Staff Software Engineer Jobs in Boston, MA
Systems Engineer Jobs in Boston, MA
All Filters
Total selected ()
No Results
No Results































