Maximum of 25 job preferences reached.
Top Senior Site Reliability Engineer Jobs in Boston, MA
Fintech • Information Technology • Payments • Sharing Economy • Financial Services • Cryptocurrency
Leads reliability, scalability, performance, and security for large-scale cloud systems. Designs AWS infrastructure with Terraform, automates CI/CD and operational workflows, establishes SLOs, manages incident response and disaster recovery, develops monitoring and observability solutions, and builds internal tools. Partners with engineering teams on architecture and reliability practices, conducts code reviews, mentors SREs, and supports compliance, vulnerability management, and security integration.
Top Skills:
Agentic ApplicationsAmazon Api GatewayAmazon AuroraAmazon CloudfrontAmazon CloudwatchAmazon DynamodbAmazon EbsAmazon Ec2Amazon EcsAmazon EfsAmazon RdsAmazon Route 53Amazon S3Amazon VpcAWSAws FargateAws LambdaAws X-RayChaos EngineeringCi/CdDastDatadogDistributed SystemsDockerEvent-Driven SystemsGitlabGitopsGrafanaIamInfrastructure As CodeJavaKubernetesLlmsMicroservicesNew RelicNode.jsOwasp Top 10PythonSastServerless ArchitecturesSplunkTerraform
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Design and implement reliable, scalable IT infrastructure; automate processes; monitor systems; resolve incidents; conduct performance and load testing; optimize cloud environments; maintain storage architecture; and lead continuous improvement. The role requires troubleshooting complex system issues, developing automation solutions, managing incidents, collaborating across teams, mentoring others, and supporting secure business operations across AWS, Google Cloud, and Microsoft Azure.
Top Skills:
AWSGoogle Cloud PlatformAzure
Reposted 18 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Information Technology • Security • Software • Cybersecurity
As a Staff Site Reliability Engineer, you'll oversee Zscaler production data center services, optimize code, and ensure cloud service availability and performance. Collaborate with cross-functional teams to improve processes and resolve escalated issues.
Top Skills:
BashDnsFirewallsGrafanaHTTPIcmpLoad BalancingNagiosOsi ModelPrometheusPythonTcp/Ip
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads architecture, modernization, optimization, and reliability initiatives for mainframe CICS, MQ, and z/OS Connect environments. Provides technical direction across development and operations teams, establishes governance and change processes, tunes performance using telemetry, resolves incidents, and develops modernization roadmaps. Collaborates with stakeholders and enterprise architects to deliver secure, scalable, high-availability solutions while evaluating automation, cloud integration, and AI technologies.
Top Skills:
AnsibleCicsCobolDevOpsIbm MqIbm Z/OsOpenshiftPythonRed Hat Ansible Automation PlatformZ/Os ConnectZlinux
17 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Security • Software • Cybersecurity • Automation
Provide technical direction for GitLab Dedicated, a managed single-tenant SaaS platform. Lead architecture and transformation across resilience, failover, tenant orchestration, change management, automation, and platform integrations. Identify systemic reliability and scalability risks, establish reusable platform patterns, strengthen service ownership, and guide cross-team technical decisions. Mentor senior engineers and advance engineering excellence across the organization.
Top Skills:
Cloud InfrastructureDevsecopsDistributed SystemsGoInfrastructure As CodeObservabilityPythonRuby
Reposted 26 Days AgoSaved
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Leads teams designing and deploying AI-driven enterprise and cloud security solutions. Responsibilities include developing machine learning systems, integrating data infrastructure and pipelines, performing advanced modeling and analysis, and building AI applications with Python, Java, and C++. The role involves client engagement, strategic problem-solving, stakeholder validation, coaching, process innovation, and operational excellence while addressing complex cybersecurity and privacy challenges.
Top Skills:
Ai SystemsAWSC++Cloud SecurityData EngineeringData PipelinesDatabricksGCPJavaMachine LearningAzurePythonScikit-LearnSnowflakeTensorFlow
Cloud • Security • Software • Cybersecurity
Performs site reliability engineering for large-scale cloud infrastructure, focusing on application and network performance, reliability, security, scalability, and capacity. Deploys highly available systems, automates cloud service deployments, monitors and troubleshoots services, analyzes logs and events, maintains SLAs, and resolves infrastructure issues. The role requires expertise in microservices, container orchestration, cloud migration, deployment automation, and multiple DevOps technologies.
Top Skills:
AzureChefCloud ComputingContainer OrchestrationGitIntellij IdeaJenkinsKubernetesLinuxMicroservicesMySQLPostgresPythonTerraformVMware
Cloud • Security • Software • Cybersecurity
Analyzes and resolves availability and performance issues in large-scale production and lab environments. Develops automation, monitoring, alerting, log analysis, debugging tools, and systems programming solutions. Collaborates with software development and engineering teams on CI/CD, platform architecture, incident resolution, and operational best practices within an agile SDLC.
Top Skills:
AlertingAnsibleAutomation ScriptingBashContinuous DeliveryContinuous IntegrationDevOpsLinuxLog AnalysisMonitoringPowershellSystems Programming
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead long-term strategy and architecture for cloud and on‑prem platform infrastructure, driving Kubernetes and multi‑cloud reliability, IaC/GitOps automation, observability, SLO/SLI/error‑budget practices, incident leadership, AI‑augmented tooling adoption, and mentorship of senior engineers to improve platform resilience and developer experience.
Top Skills:
Amazon Elastic Kubernetes Service (Eks)AutoscalingAWSCapacity PlanningCi/CdGitopsGoGoogle Cloud PlatformGoogle Kubernetes Engine (Gke)Identity And Access ManagementInfrastructure As CodeKubernetesLinuxNetworkingObservabilityOperatorsPulumiPythonRke2StorageTerraform
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Build and maintain the Zero Touch Platform: develop Temporal workflows and Go services to automate cloud operations, own production systems (SLOs, on-call, incidents), debug cloud-native distributed systems, create CI/CD and IaC automation, and produce documentation to enable self-service across engineering teams.
Top Skills:
AksApmAWSAws CloudformationAzureC#Ci/CdEksGoJavaKubernetesLoggingMetricsPythonTemporalTerraform
Cloud • Software
Support Salesforce GovCloud reliability and high availability in a 24/7 operations environment. Responsibilities include incident response, root cause analysis, infrastructure administration, security and compliance alignment, automation, monitoring, preventative remediation, and workflow improvement. The role requires collaboration across engineering teams, participation in post-incident reviews, and support for AWS/C2S infrastructure and customer-facing services.
Top Skills:
Ai ToolsAWSAws CliAws SdksBambooBsdC2SChefGoJavaJenkinsKubernetesLinuxPuppetPythonRed Hat Enterprise LinuxSolarisSpinnakerTcp/IpUnix
Fintech • Payments • Financial Services
Leads the design and operation of highly available cloud systems, primarily on AWS, using Terraform, GitLab CI/CD, containers, and observability tools. Responsibilities include reliability engineering, incident response, disaster recovery, automation, monitoring, security integration, vulnerability management, and cost optimization. The role builds internal tools, guides architecture, mentors SRE engineers, documents operational processes, and partners with software engineering teams to improve production reliability and compliance.
Top Skills:
Agentic ApplicationsAmazon Api GatewayAmazon AuroraAmazon CloudfrontAmazon CloudwatchAmazon DynamodbAmazon EbsAmazon Ec2Amazon EcsAmazon EfsAmazon RdsAmazon Route 53Amazon S3Amazon VpcAWSAws FargateAws LambdaAws X-RayDastDatadogDockerGitlabGrafanaIamJavaKubernetesLlmsNew RelicNode.jsOwasp Top 10PythonSastSplunkTerraform
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3 • Infrastructure as a Service (IaaS)
Manage AWS and GCP cloud environments, scale infrastructure globally, shape technical architecture, and build reliable CI/CD pipelines. Automate security and compliance controls, improve infrastructure performance, and develop systems interacting with smart contracts across multiple blockchains. The role requires Terraform, shell scripting, GitHub Actions, Docker, production cloud operations, and observability experience, with Kubernetes, networking, and fintech compliance knowledge preferred.
Top Skills:
AWSCi/CdDockerFirewallsGCPGithub ActionsGkeGoHelmInfrastructure As CodeKubernetesLoad BalancersMtlsNode.jsPciShellSsl/TlsTerraformTypescriptVpcZero Trust
Cloud • Security • Software • Cybersecurity
Ensure reliability, scalability, and usability of network infrastructure for Akamai Connected Cloud. Define requirements and SLOs, build automation and CI/CD pipelines, collaborate with dev/QA to improve code and stability, troubleshoot complex network issues (on-call), and mentor teammates while driving architectural standards.
Top Skills:
AnsibleArgocdBashBirdChefFrrGithub ActionsGoGobgpJenkinsLinux NetworkingPuppetPythonSalt Stack
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills:
BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
Cloud • Information Technology • Internet of Things • Professional Services • Software
Deploys, operates, and maintains resilient AWS and Kubernetes infrastructure and microservices for the Webex for Government environment. Monitors service health, troubleshoots production incidents, automates infrastructure provisioning, and supports 24x7 operations. The role strengthens security and FedRAMP compliance through IAM, logging, hardening, vulnerability remediation, and continuous monitoring. It also partners with engineering, security, and compliance teams to improve reliability, document procedures, create runbooks, and conduct incident and root-cause reviews.
Top Skills:
Amazon CloudwatchAWSAws CloudtrailBashCi/CdDockerElastic StackFedrampGitGitlabGoGrafanaJavaJenkinsKubernetesLinuxMicroservicesPrometheusPythonSplunkTerraform
Software
Own the operational health, observability, reliability, and failover of AI model providers and endpoints. Build SLO monitoring, canaries, quality regression detection, automated provider-operations tooling, and load-testing systems. Lead incident response, communicate with external providers, produce reliability scorecards, and drive postmortems and corrective actions. Support capacity planning and launch readiness for high-volume LLM inference traffic.
Top Skills:
ClickhouseCloudflare WorkersGCPPostgresPythonTypescriptVercel
eCommerce
Own the reliability, availability, security, and observability of Tradeweb’s global AWS platform. Responsibilities include infrastructure-as-code automation, monitoring, SLO development, incident triage and resolution, performance analysis, architecture collaboration, and regular on-call support. The role requires cloud-native engineering, scripting, networking, Linux/Unix, security, and reliability expertise while contributing to an Agile engineering organization.
Top Skills:
Amazon EksAmazon SmsAmazon SnsArgocdAWSAws LambdaGitsecopsKubernetesKustomizeLgtmLinuxPulumiPythonUnix
Big Data • Software
Owns production reliability for a B2B SaaS platform, including SLOs, error budgets, observability, alerting, incident response, recovery exercises, performance and capacity engineering, load testing, disaster recovery, and production readiness. The role writes automation and production code, leads incidents, coaches engineering teams, and improves reliability across Kubernetes, Azure, databases, messaging, and service infrastructure.
Top Skills:
.NetAksAsp.NetAWSAzureAzure Service BusC#Dora MetricsFluxGatlingGitopsGoGrafanaIncident.IoIstioJmeterK6KafkaKubernetesLokiMongoDBNist 800-171PagerdutyPrometheusPythonSentrySignalrSoc 2TempoTemporalTypescript
Software
Own production reliability and infrastructure operations for the GrayKey cloud platform. Responsibilities include AWS monitoring, incident response, vulnerability remediation, networking, IAM and Okta administration, EKS and Argo CD operations, backups, recovery testing, Terraform changes, cost optimization, automation, documentation, and compliance support. The role requires independent work during Pacific and Mountain time-zone hours and involves less than 5% travel.
Top Skills:
Active DirectoryAmazon EksAnsibleArgo CdAWSAzure AdBashDatadogDnsDockerEc2Entra IdFortigateGitlab Ci/CdHelmIamKubernetesLambdaLinuxOidcOktaPostgresPythonRdsRedisSAMLTcp/IpTerraformTlsVpcVpn
Artificial Intelligence • Information Technology • Consulting
The Linux Systems Administrator will maintain and troubleshoot Linux systems, support network services, and work on systems integration while collaborating with infrastructure teams.
Top Skills:
DhcpDnsLinuxNtpPython
Consumer Web • Gaming • Mobile • News + Entertainment • Software
Leads reliability practices across Infrastructure Engineering by defining Service Level Objectives, Service Level Indicators, error budgets, observability standards, and reporting. Partners with engineering teams to connect infrastructure reliability to application and platform outcomes, analyzes distributed-system failure modes, builds Datadog dashboards and monitoring, and influences reliability strategy through technical guidance, mentoring, and executive communication.
Top Skills:
Amazon Web ServicesDatadogDistributed SystemsError BudgetsKubernetesObservabilityService Level IndicatorsService Level ObjectivesSre
Consumer Web • Gaming • Mobile • News + Entertainment • Software
Leads the strategy, architecture, and operation of Kubernetes-based cloud and on-premise infrastructure. Drives reliability engineering, Infrastructure as Code, GitOps, observability, automation, incident response, capacity planning, and cost optimization. Leads cross-functional platform initiatives, establishes SLOs and error budgets, mentors engineers, and advances AI-enabled engineering practices across the organization.
Top Skills:
Amazon Elastic Kubernetes ServiceAWSCi/CdDistributed SystemsGitopsGoGoogle Cloud PlatformGoogle Kubernetes EngineInfrastructure As CodeKubernetesLinuxObservabilityPulumiPythonRke2Terraform
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead reliability, scalability, and operational excellence of large-scale database platforms across cloud and on-prem. Build automation-first database infrastructure (Kubernetes operators, IaC, GitOps), drive monitoring/SLOs, incident leadership, performance and cost optimization, and partner with application teams on safe schema/migration practices. Mentor engineers and evaluate AI-assisted workflows to improve productivity and reliability.
Top Skills:
AerospikeArgocdAuroraClaudeCloud SqlCursorDatabase OperatorsEksFluxcdGithub CopilotGitopsGkeGoKubernetesMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform
Information Technology
Leads SRE and cloud operations for highly available, scalable platforms and AI-powered solutions. Designs AWS infrastructure, Kubernetes environments, CI/CD pipelines, Infrastructure as Code, observability, automation, and incident response processes. Develops AI-driven operational capabilities, supports MLOps and model lifecycle management, establishes reliability metrics and SLOs, and mentors engineering teams. Partners with software, data, machine learning, security, and product teams to improve platform performance, resilience, and operational excellence.
Top Skills:
Amazon SagemakerAWSAws BedrockAzure DevopsAzure OpenaiCi/CdCloudFormationCloudwatchDatadogDockerEc2EcsEksElk StackGithub ActionsGitlab Ci/CdGrafanaIamInfrastructure As CodeJenkinsKubernetesLambdaLangchainMlopsNvidia AiOpenai ApisPrometheusPythonRdsS3ShellSplunkTerraformVpc
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Boston Companies Hiring Senior Site Reliability Engineers
See AllPopular Job Searches
All Filters
Total selected ()
No Results
No Results





























