Maximum of 25 job preferences reached.
Top Hybrid DevOps & Platform Engineering Jobs in Boston, MA
AdTech • eCommerce • Food • Marketing Tech • Retail
The Cloud Reliability Engineer III designs and implements cloud services, automates operations, and collaborates with teams to enhance reliability and developer experience.
Top Skills:
AdoAnsibleArmAzcliAzure AutomationAzure FunctionsAzure IaasAzure Log AnalyticsAzure MonitorAzure PaasAzure Security CenterContainersGitKubernetesOpenshiftPowershellPythonTerraform
AdTech • eCommerce • Food • Marketing Tech • Retail
The Cloud Reliability Engineer III designs and implements cloud services, automates operations, enhances reliability, and collaborates on critical IT escalations.
Top Skills:
AnsibleAzcliAzure AutomationAzure FunctionsAzure GovernanceAzure IaasAzure IdentityAzure MonitoringAzure PaasAzure SecurityAzure Virtual MachinesAzure Virtual NetworkContainersGitKubernetesOpenshiftPowershellPythonTerraform
AdTech • eCommerce • Food • Marketing Tech • Retail
The Azure Cloud Engineer III designs cloud services, automates operations, enhances reliability, and troubleshoots incidents while collaborating with teams on projects.
Top Skills:
AnsibleAzure AutomationAzure DevopsAzure FunctionsAzure IaasAzure MonitorAzure PaasContainersGitKubernetesLog AnalyticsOpenshiftPowershellPythonTerraform
Consumer Web • eCommerce • Software
Design and deliver developer tooling, core frameworks, and templates to improve engineering productivity. Advance agentic AI tooling (Claude Code, plugins, MCP), strengthen authentication/identity (Keycloak, OAuth2/OIDC), and build/operate services on Java, Spring Boot, AWS, and Kubernetes. Define standards, improve reliability, and partner across teams to drive adoption.
Top Skills:
AWSClaude CodeJavaJavaScriptKeycloakKubernetesMcpNode.jsOauth2OidcReactRemixSpring Boot
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead the design, automation, and scaling of global compute infrastructure across data centers, cloud, and on-prem. Operate GitOps with Rancher Fleet/Flux/Helm, build self-healing tooling, own cluster autoscaling and capacity strategy, define SLOs using Datadog, and participate in on-call rotation while mentoring peers.
Top Skills:
AWSContainerdDatadogDockerFluxGCPGitopsGoHelmHpaInfrastructure As Code (Iac)KarpenterKedaKubernetesLinuxNutanixPythonRancher FleetVsphere
Reposted 18 Hours AgoSaved
Fintech • Machine Learning • Payments • Software • Financial Services
Lead design, development, deployment, and support of foundational AI systems including foundation model training, LLM inference, similarity search, guardrails, evaluation, and observability. Optimize large-scale AI performance (cost, latency, throughput), partner cross-functionally, and contribute to technical vision and roadmap.
Top Skills:
Aws UltraclustersC#C++GoHuggingfaceJavaLlm InferenceNemo GuardrailsPythonPyTorchScalaSimilarity SearchVectordbs
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Reposted YesterdaySaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Lead design and implementation of LLM observability features: prototype and scale product capabilities for tracing, evaluating, and debugging generative AI systems. Work cross-functionally to influence architecture, mentor engineers, prioritize customer pain points, and drive product and engineering decisions for reliable, high-performance AI observability.
Top Skills:
Distributed SystemsGenerative AiInference PipelinesLarge Language Models (Llms)Observability Tools/PlatformsPrompt EngineeringScalable Backend Architectures
Cloud • Information Technology • Security • Software • Cybersecurity
As a Staff Site Reliability Engineer, you'll oversee Zscaler production data center services, optimize code, and ensure cloud service availability and performance. Collaborate with cross-functional teams to improve processes and resolve escalated issues.
Top Skills:
BashDnsFirewallsGrafanaHTTPIcmpLoad BalancingNagiosOsi ModelPrometheusPythonTcp/Ip
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Design, implement, and operate cloud-native HPC and ML/AI infrastructure (AWS/GCP). Lead containerization, IaC (Terraform/CloudFormation), orchestration (Kubernetes/EKS/GKE), monitoring, and migration. Collaborate with HPC specialists to automate deployments, optimize performance, and maintain platform reliability, security, and cost efficiency.
Top Skills:
ApptainerAWSAws ParallelclusterCloudFormationCloudwatchDockerEksGCPGkeGoogle Cloud Cluster ToolkitGrafanaKubernetesLinuxNvidia Gpu ComputingOpen On DemandPrometheusSingularitySlurmTerraform
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Lead enterprise-scale cloud modernization and migration programs to AWS, design end-to-end architectures (including AI agent integration), define AI-native engineering standards, mentor engineering teams, implement IaC and CI/CD, enforce security and governance, and deliver measurable business value through technical strategy and execution.
Top Skills:
AngularApi GatewayAWSAws CdkBedrock Agent SdkCi/CdClaude Agent SdkClaude CodeCloudFormationCloudfrontCloudtrailCloudwatchCodepipelineCodexCognitoCursorDmsDockerDynamoDBEc2EcsEksGithub ActionsGithub CopilotIamJavaJenkinsKiroKubernetesLambdaLangchainMigration HubNode.jsPythonRdsReactS3SctTerraformVueX-Ray
Artificial Intelligence • Fintech • Greentech • Sales • Software • Travel • Hospitality
Design, build, and maintain scalable AWS infrastructure and developer tooling. Enable teams to own systems, support CI/CD, observability, containers, on-call/incident management, and collaborate with SRE/security to ensure resilient, secure services.
Top Skills:
AWSBashCircleCICloudwatchDatadogDockerEc2EcsGithub ActionsGoIamPulumiPythonRdsS3TerraformVpc
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top hybrid Companies in Boston, MA Hiring Engineering Roles
See AllAll Filters
Total selected ()
No Results
No Results





















