Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in Boston, MA
Cloud • Security • Software • Cybersecurity
Ensure reliability, scalability, and usability of network infrastructure for Akamai Connected Cloud. Define requirements and SLOs, build automation and CI/CD pipelines, collaborate with dev/QA to improve code and stability, troubleshoot complex network issues (on-call), and mentor teammates while driving architectural standards.
Top Skills:
AnsibleArgocdBashBirdChefFrrGithub ActionsGoGobgpJenkinsLinux NetworkingPuppetPythonSalt Stack
Cloud • Security • Software • Cybersecurity
Lead and mentor SRE teams; partner with engineering, operations and product; apply statistical analysis and networking expertise to diagnose performance and reliability issues; define and implement data feeds; influence technical decisions and investments; build tooling to automate analytical workflows and increase platform reliability.
Top Skills:
CCloudDistributed SystemsDnsEdgeHTTPJavaPerlPythonRSQLTcpTls
Software
The role involves managing compute infrastructure for decentralized applications, requiring critical thinking, documentation skills, and experience in Kubernetes and blockchain management.
Top Skills:
BlockchainGitopsInfrastructure-As-CodeKubernetesProgramming Languages
Reposted 21 Hours AgoSaved
Other • Social Impact
As a Senior Site Reliability Engineer, you will design, develop, and maintain reliable infrastructure for Wikimedia's API services, ensuring performance and availability while driving reliability engineering practices and improving developer experience.
Top Skills:
AnsibleArgocdAWSAzureGCPGitlabGoKubernetesOpentelemetryPrometheusPythonTerraform
Social Media • Software
Design, implement, and operate infrastructure for a federated social network. Own reliability, availability, observability, incident response, deployments, capacity planning, and cost management. Build automation and tooling, scale bare-metal and cloud systems for millions of users, lead incident reviews, mentor engineers, and manage vendor relationships to ensure operational excellence.
Top Skills:
Bare-MetalCapacity PlanningCloud ServicesColocationDatabasesDeployment And Rollback SystemsGoIncident ResponseKubernetesLinuxMonitoringNetworkingObservability SystemsProduction AutomationStorage
Healthtech • Software
Design, automate, and maintain scalable infrastructure and SRE tooling. Manage Kubernetes clusters, CI/CD, monitoring, and incident response. Improve processes, reduce toil via automation, and collaborate with engineering and data teams to support domestic and international workloads.
Top Skills:
AWSAzureContainerdDnsDockerFirewallsGCPGoGrpcHelmKubernetesLinuxLoad BalancingPrometheusPythonRoutingShell ScriptingTcp/IpUdp
Financial Services
Prototype, write, test, document, and deploy release automation across environments. Build and maintain pipelines, collaborate with engineers and product teams, troubleshoot issues, participate in on-call rotation, and improve software delivery, configuration, monitoring, and operations.
Top Skills:
AnsibleBashDockerGitlabJenkinsKubernetesMssqlPostgresPowershellPythonRedisTeamcity
Software • Consulting
Lead 24x7 application support for external web applications: manage incidents, perform RCA, implement preventative fixes, build monitoring/alerts, expand Splunk functionality, create dashboards, collaborate with development and platform teams, and participate in on-call rotation.
Top Skills:
ApmAppdynamicsAWSDatadogGrafanaKubernetesLinuxMulesoftOpenshiftOpentelemetryPostmanPythonRumSeleniumServicenowShell ScriptingSplunk (Spl)Splunk CloudSplunk Observability CloudSplunk Synthetics
Artificial Intelligence • Software • Generative AI • Automation
Lead design, build, and operation of scalable, fault-tolerant cloud infrastructure. Define SLOs/SLAs, improve observability and incident response, own CI/CD and deployment automation, partner with engineering teams on reliability, capacity planning, performance benchmarking, cost optimization, and security for an AI platform.
Top Skills:
AWSAzureBashCi/CdDatadogEbpfGCPGoGpuGrafanaIstioKubernetesLinkerdOpentelemetryPrometheusPulumiPythonTerraform
Software
Lead SRE to define SRE strategy, architecture, and roadmap; design and operate containerized, compliant cloud environments; build observability, incident management, automation, and developer platform capabilities; mentor SRE team and collaborate with security, compliance, and product teams to ensure reliability at scale.
Top Skills:
AWSAws MarketplaceAzureAzure MarketplaceGCPGoogle Cloud MarketplaceGrafanaKubernetesPrometheusTerraform
Artificial Intelligence • Other • Sales • Software
The role involves designing and advancing infrastructure for the engineering team, ensuring the reliability of Kubernetes clusters, automating operations, and building machine learning infrastructure.
Top Skills:
ArgoAWSAzureCloudFormationFluxGithub ActionsGoGCPKubernetesPostgresPythonTerraform
Agency • Information Technology
Lead SRE role designing and maintaining CI/CD pipelines (GitHub Actions), containerized deployments (Docker, Kubernetes, AKS, Helm), web/mobile app releases, observability, automated testing, and DevOps best practices across cloud environments with cross-functional collaboration and regulatory compliance.
Top Skills:
AksAndroidAzure Application InsightsAzure Log AnalyticsAzure MonitorBashBranchingDockerDocker ComposeGitGit HooksGithub ActionsGoogle PlayHelmHerokuiOSIos App StoreJavaKubernetesNpmPowershellPull RequestsPythonSonarqubeVeracodeVercel
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Artificial Intelligence • Insurance • Software • Automation
The Staff Site Reliability Engineer will build and scale infrastructure for Assured's platform, automate delivery, enhance observability, and lead mentoring initiatives.
Top Skills:
AWSKubernetesPostgresTerraform
Healthtech • Social Impact • Software
Own the operational lifecycle of cloud-native data infrastructure: design and automate reliable deployments, observability, incident response, SLIs/SLOs, autoscaling and IaC, and improve platform efficiency and data freshness across GKE and Cloud Run.
Top Skills:
BashBigQueryCloud BuildCloud MonitoringCloud RunDatadogDockerGCPGithub ActionsGkeGoGrafanaJIRAKubernetesPrometheusPulumiPythonSentrySlackSnykSonarqubeTerraform
Software
Own and improve platform performance, reliability, and deployment automation. Manage cloud infrastructure, implement IaC, monitor systems with observability tools, provide operational support for distributed applications, and integrate production learnings into development workflows.
Top Skills:
Aiops ToolingAws Elastic ContainersAws RdsAws S3Claude CodeClaude CoworkDatadogHarness EngineeringInfrastructure As CodeKubernetesLlmsPrompt EngineeringRigorSplunk
Software
Senior SRE responsible for production reliability of Grafana Cloud databases (Mimir, Loki, Tempo, Pyroscope). Partner with product squads, define per-tenant SLOs, automate reliability practices, lead incident response/on-call, reduce toil, improve alerting, and influence design for scalability and operability.
Top Skills:
AWSAzureGCPGoGrafana CloudHelmJavaJsonnetKubernetesLinuxLokiMimirPyroscopePythonTempoTerraform
Artificial Intelligence • Healthtech • Software • Telehealth
Design, deploy, and maintain AWS-hosted Kubernetes (EKS) infrastructure; build automation and AI-assisted runbooks; provide observability and incident response; ensure HIPAA-compliant, high-availability platform operations and mentor engineering teams.
Top Skills:
AWSBashDatadogEc2EksGitGithub ActionsGoHelmKubernetesPythonRdsS3Terraform
Other
The Senior Site Reliability Engineer at Juul Labs ensures operational stability and performance of hybrid cloud infrastructure, leads automation, and handles critical incidents.
Top Skills:
AWSBashCloudFormationGCPNutanixPowershellPythonTerraform
Cloud • Information Technology • Internet of Things • Professional Services • Software
Lead design, build, and evolve developer infrastructure and CI platforms for Meraki cloud teams. Guide complex troubleshooting, mentor engineers, drive operational excellence, define roadmaps with leadership, and champion sustainable on-call practices while supporting large-scale distributed systems and automation across developer environments.
Top Skills:
Artifact ManagementBare MetalBuild ToolsCiCi PlatformsCode ReviewConfiguration-As-CodeContainer OrchestrationContainerizationInfrastructure AutomationManaged Cloud ServicesPythonRubyUnix/Linux
Digital Media • Social Media • Software • Sports
Lead the technical architecture and execution of migration to AWS, drive developer enablement, and automate infrastructure using code-first principles.
Top Skills:
Aws EksDatadogGithub ActionsGoIstioK6KubernetesNode.jsTerraform
Computer Vision • Machine Learning • Software
As a Site Reliability Engineer, ensure the reliability, performance, and scalability of Ditto's cloud infrastructure by developing observability solutions, leading incident management, and collaborating with product engineering teams.
Top Skills:
AWSAzureCDatadogGCPGoGrafanaHelmJavaKubernetesPrometheusRustTerraform
Artificial Intelligence • Fintech • Machine Learning • Natural Language Processing • Business Intelligence
Lead architecture and implementation of reliability platforms and SRE practices for a production SaaS. Build self-service reliability tooling, drive AIOps automation, advance observability (monitoring, tracing, profiling), lead incident response and postmortems, mentor engineers, and embed production readiness across teams to achieve 99.99% uptime.
Top Skills:
AWSAzureContinuous ProfilingDatadogDnsElkGCPGoGrafanaHttp/SKubernetesLoad BalancingOpentelemetryPrometheusPythonTcp/Ip
Database • Analytics
This role involves ensuring the reliability and performance of ClickHouse's cloud infrastructure, collaborating with engineering teams, incident management, and driving continuous improvement in service availability.
Top Skills:
AnsibleAWSAzureClickhouseDocker SwarmGoGoogle Cloud PlatformKubernetesPuppetPythonTerraform
Software • Financial Services
Ensure platform reliability, performance, and availability by implementing observability, automating infrastructure, participating in on-call rotations and post-mortems, partnering with Product and Engineering, designing scalable architectures, mentoring teammates, and integrating Dynatrace with Azure DevOps and Jira while supporting compliance (SOC/FedRAMP).
Top Skills:
.NetAksAlpineAnsibleAppinsightsArm TemplatesAWSAzure DevopsBashBicepC#ChefCloudFormationDatadogDebianDynatraceEksGCPGitGitGksGrafanaHelmJIRAKubernetesLog AnalyticsAzureNew RelicOnestream SoftwareOpenshiftPowershellPowershell DscPrometheusPuppetPythonRest ApisSQLTerraformUbuntu
Fintech • Information Technology
As a Site Reliability Engineer at Alpaca, you will ensure system reliability and performance, troubleshoot issues, and collaborate with teams to design scalable features.
Top Skills:
GoGormLinuxPgxPostgresPrometheusSqlc
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Boston, MA Companies Hiring SRE Engineers
See AllPopular Boston, MA Engineering Job Searches
Engineering Jobs in Boston, MA
Software Engineer Jobs in Boston, MA
Android Developer Jobs in Boston, MA
C# Jobs in Boston, MA
C++ Jobs in Boston, MA
DevOps Jobs in Boston, MA
Front End Developer Jobs in Boston, MA
Golang Jobs in Boston, MA
Hardware Engineer Jobs in Boston, MA
iOS Developer Jobs in Boston, MA
Java Developer Jobs in Boston, MA
Javascript Jobs in Boston, MA
Linux Jobs in Boston, MA
Engineering Manager Jobs in Boston, MA
.NET Developer Jobs in Boston, MA
PHP Developer Jobs in Boston, MA
Python Jobs in Boston, MA
QA Jobs in Boston, MA
Ruby Jobs in Boston, MA
Salesforce Developer Jobs in Boston, MA
Scala Jobs in Boston, MA
Application Engineer Jobs in Boston, MA
Automation Engineer Jobs in Boston, MA
AWS Engineer Jobs in Boston, MA
Backend Engineer Jobs in Boston, MA
Cloud Engineer Jobs in Boston, MA
Controls Engineer Jobs in Boston, MA
CTO Jobs in Boston, MA
Design Engineer Jobs in Boston, MA
DevOps Engineer Jobs in Boston, MA
Director of Engineering Jobs in Boston, MA
Electrical Engineering Jobs in Boston, MA
Embedded Software Engineer Jobs in Boston, MA
Full-Stack Engineer Jobs in Boston, MA
Game Engineer Jobs in Boston, MA
Infrastructure Engineer Jobs in Boston, MA
Manufacturing Engineer Jobs in Boston, MA
Mechanical Design Engineer Jobs in Boston, MA
Mechanical Engineer Jobs in Boston, MA
Mechatronics Engineering Jobs in Boston, MA
Network Engineer Jobs in Boston, MA
Platform Engineer Jobs in Boston, MA
Principal Engineer Jobs in Boston, MA
Principal Software Engineer Jobs in Boston, MA
Process Engineer Jobs in Boston, MA
Product Engineer Jobs in Boston, MA
Project Engineer Jobs in Boston, MA
QA Engineer Jobs in Boston, MA
Robotics Engineer Jobs in Boston, MA
Security Engineer Jobs in Boston, MA
Software Engineering Manager Jobs in Boston, MA
Software Test Engineer Jobs in Boston, MA
Solutions Architect Jobs in Boston, MA
Solutions Engineer Jobs in Boston, MA
SRE Engineer Jobs in Boston, MA
Staff Software Engineer Jobs in Boston, MA
Systems Engineer Jobs in Boston, MA
All Filters
Total selected ()
No Results
No Results



































