Top Hybrid DevOps & Platform Engineering Jobs in Boston, MA

YesterdaySaved
Easy Apply
Hybrid
Boston, MA, USA
Easy Apply
244K-305K Annually
Senior level
244K-305K Annually
Senior level
Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Build and maintain Bazel-based build, test, and packaging tools for a large monorepo. Improve CI performance and cost efficiency, run complex migrations, keep pipelines reliable, and contribute upstream to the Bazel ecosystem. Own projects end-to-end and collaborate with developers to improve tool usability and developer productivity.
Top Skills: BazelCiGoJavaMonorepoPythonRustTypescript
Reposted YesterdaySaved
Hybrid
Quincy, MA, USA
125K-188K Annually
Senior level
125K-188K Annually
Senior level
AdTech • eCommerce • Food • Marketing Tech • Retail
The Cloud Reliability Engineer III designs and implements cloud services, automates operations, and collaborates with teams to enhance reliability and developer experience.
Top Skills: AdoAnsibleArmAzcliAzure AutomationAzure FunctionsAzure IaasAzure Log AnalyticsAzure MonitorAzure PaasAzure Security CenterContainersGitKubernetesOpenshiftPowershellPythonTerraform
Reposted YesterdaySaved
Hybrid
Boston, MA, USA
140K-176K Annually
Senior level
140K-176K Annually
Senior level
Consumer Web • eCommerce • Software
Design and deliver developer tooling, core frameworks, and templates to improve engineering productivity. Advance agentic AI tooling (Claude Code, plugins, MCP), strengthen authentication/identity (Keycloak, OAuth2/OIDC), and build/operate services on Java, Spring Boot, AWS, and Kubernetes. Define standards, improve reliability, and partner across teams to drive adoption.
Top Skills: AWSClaude CodeJavaJavaScriptKeycloakKubernetesMcpNode.jsOauth2OidcReactRemixSpring Boot
Reposted YesterdaySaved
Hybrid
Boston, MA, USA
128K-160K Annually
Senior level
128K-160K Annually
Senior level
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead the design, automation, and scaling of global compute infrastructure across data centers, cloud, and on-prem. Operate GitOps with Rancher Fleet/Flux/Helm, build self-healing tooling, own cluster autoscaling and capacity strategy, define SLOs using Datadog, and participate in on-call rotation while mentoring peers.
Top Skills: AWSContainerdDatadogDockerFluxGCPGitopsGoHelmHpaInfrastructure As Code (Iac)KarpenterKedaKubernetesLinuxNutanixPythonRancher FleetVsphere
Reposted YesterdaySaved
Hybrid
Cambridge, MA, USA
230K-286K Annually
Senior level
230K-286K Annually
Senior level
Fintech • Machine Learning • Payments • Software • Financial Services
Lead design, development, deployment, and support of foundational AI systems including foundation model training, LLM inference, similarity search, guardrails, evaluation, and observability. Optimize large-scale AI performance (cost, latency, throughput), partner cross-functionally, and contribute to technical vision and roadmap.
Top Skills: Aws UltraclustersC#C++GoHuggingfaceJavaLlm InferenceNemo GuardrailsPythonPyTorchScalaSimilarity SearchVectordbs
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted 2 Days AgoSaved
Easy Apply
Remote or Hybrid
2 Locations
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills: AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Reposted 3 Days AgoSaved
Easy Apply
Remote or Hybrid
Boston, MA, USA
Easy Apply
119K-170K Annually
Senior level
119K-170K Annually
Senior level
Cloud • Information Technology • Security • Software • Cybersecurity
As a Staff Site Reliability Engineer, you'll oversee Zscaler production data center services, optimize code, and ensure cloud service availability and performance. Collaborate with cross-functional teams to improve processes and resolve escalated issues.
Top Skills: BashDnsFirewallsGrafanaHTTPIcmpLoad BalancingNagiosOsi ModelPrometheusPythonTcp/Ip
3 Days AgoSaved
Hybrid
Cambridge, MA, USA
124K-207K Annually
Senior level
124K-207K Annually
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Design, implement, and operate cloud-native HPC and ML/AI infrastructure (AWS/GCP). Lead containerization, IaC (Terraform/CloudFormation), orchestration (Kubernetes/EKS/GKE), monitoring, and migration. Collaborate with HPC specialists to automate deployments, optimize performance, and maintain platform reliability, security, and cost efficiency.
Top Skills: ApptainerAWSAws ParallelclusterCloudFormationCloudwatchDockerEksGCPGkeGoogle Cloud Cluster ToolkitGrafanaKubernetesLinuxNvidia Gpu ComputingOpen On DemandPrometheusSingularitySlurmTerraform
3 Days AgoSaved
Hybrid
Boston, MA, USA
124K-280K Annually
Senior level
124K-280K Annually
Senior level
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Lead enterprise-scale cloud modernization and migration programs to AWS, design end-to-end architectures (including AI agent integration), define AI-native engineering standards, mentor engineering teams, implement IaC and CI/CD, enforce security and governance, and deliver measurable business value through technical strategy and execution.
Top Skills: AngularApi GatewayAWSAws CdkBedrock Agent SdkCi/CdClaude Agent SdkClaude CodeCloudFormationCloudfrontCloudtrailCloudwatchCodepipelineCodexCognitoCursorDmsDockerDynamoDBEc2EcsEksGithub ActionsGithub CopilotIamJavaJenkinsKiroKubernetesLambdaLangchainMigration HubNode.jsPythonRdsReactS3SctTerraformVueX-Ray
4 Days AgoSaved
Hybrid
Boston, MA, USA
115K-135K Annually
Mid level
115K-135K Annually
Mid level
Artificial Intelligence • Fintech • Greentech • Sales • Software • Travel • Hospitality
Design, build, and maintain scalable AWS infrastructure and developer tooling. Enable teams to own systems, support CI/CD, observability, containers, on-call/incident management, and collaborate with SRE/security to ensure resilient, secure services.
Top Skills: AWSBashCircleCICloudwatchDatadogDockerEc2EcsGithub ActionsGoIamPulumiPythonRdsS3TerraformVpc
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account