Maximum of 25 job preferences reached.
Top Hybrid DevOps & Platform Engineering Jobs in Boston, MA
Reposted 16 Hours AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Lead design and implementation of LLM observability features: prototype and scale product capabilities for tracing, evaluating, and debugging generative AI systems. Work cross-functionally to influence architecture, mentor engineers, prioritize customer pain points, and drive product and engineering decisions for reliable, high-performance AI observability.
Top Skills:
Distributed SystemsGenerative AiInference PipelinesLarge Language Models (Llms)Observability Tools/PlatformsPrompt EngineeringScalable Backend Architectures
Reposted 17 Hours AgoSaved
Easy Apply
Easy Apply
Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Partner with Field and Product teams to design and implement LLM observability architectures, build proofs-of-concept, produce technical collateral, and advise customers to drive adoption and product feedback.
Top Skills:
DatadogJavaScriptLlmLoggingMetricsObservabilityPythonTracingTypescript
AdTech • eCommerce • Food • Marketing Tech • Retail
Lead and optimize enterprise platform operations including cluster upgrades, resource tuning, automation via IaC/GitOps, and secrets/RBAC controls. Manage VMware/Nutanix/Red Hat environments, monitor SLOs/SLIs, resolve advanced incidents, mentor junior staff, and maintain documentation, capacity planning, and multi-tenant governance.
Top Skills:
AnsibleApi GatewayAutoscalingCmdbEnterprise StorageGitopsInfrastructure As Code (Iac)IngressKubernetesLinuxNutanixObservabilityOpenshiftRbacRed HatRed Hat SatelliteSecrets ManagementService MeshSlo/SliVMwareVmware Vsphere
AdTech • eCommerce • Food • Marketing Tech • Retail
Lead and optimize enterprise platform operations including cluster upgrades, autoscaling, resource tuning, automation with IaC/GitOps, secrets and RBAC controls, and advanced incident response. Manage VMware, Nutanix, Red Hat environments, monitor SLOs/SLIs, maintain documentation/SOPs, mentor junior engineers, and drive lifecycle and multi-tenant governance.
Top Skills:
AnsibleGitopsInfrastructure As Code (Iac)KubernetesNutanixOpenshiftRbacRed Hat Enterprise LinuxRed Hat SatelliteSecrets ManagementVMware
Cloud • Information Technology • Security • Software • Cybersecurity
As a Staff Site Reliability Engineer, you'll oversee Zscaler production data center services, optimize code, and ensure cloud service availability and performance. Collaborate with cross-functional teams to improve processes and resolve escalated issues.
Top Skills:
BashDnsFirewallsGrafanaHTTPIcmpLoad BalancingNagiosOsi ModelPrometheusPythonTcp/Ip
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Design, implement, and operate cloud-native HPC and ML/AI infrastructure (AWS/GCP). Lead containerization, IaC (Terraform/CloudFormation), orchestration (Kubernetes/EKS/GKE), monitoring, and migration. Collaborate with HPC specialists to automate deployments, optimize performance, and maintain platform reliability, security, and cost efficiency.
Top Skills:
ApptainerAWSAws ParallelclusterCloudFormationCloudwatchDockerEksGCPGkeGoogle Cloud Cluster ToolkitGrafanaKubernetesLinuxNvidia Gpu ComputingOpen On DemandPrometheusSingularitySlurmTerraform
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Lead enterprise-scale cloud modernization and migration programs to AWS, design end-to-end architectures (including AI agent integration), define AI-native engineering standards, mentor engineering teams, implement IaC and CI/CD, enforce security and governance, and deliver measurable business value through technical strategy and execution.
Top Skills:
AngularApi GatewayAWSAws CdkBedrock Agent SdkCi/CdClaude Agent SdkClaude CodeCloudFormationCloudfrontCloudtrailCloudwatchCodepipelineCodexCognitoCursorDmsDockerDynamoDBEc2EcsEksGithub ActionsGithub CopilotIamJavaJenkinsKiroKubernetesLambdaLangchainMigration HubNode.jsPythonRdsReactS3SctTerraformVueX-Ray
Consumer Web • eCommerce • Software
Lead design, implementation, and operation of cloud platform services to enable self-service IaC, secrets management (Vault), CI/CD maturity (GitHub Actions, Concourse), AWS provisioning, and AI infrastructure governance (Amazon Bedrock). Own delivery of platform capabilities, enforce policy-as-code (Sentinel/Semgrep), mentor engineers, participate in on-call, and collaborate cross-team to improve reliability, security, and developer experience.
Top Skills:
Amazon BedrockAmazon QApproleAws Cost ExplorerAws IamCircleCIClaude CodeCloudzeroConcourseEc2EksElasticacheGithub ActionsGithub ConnectGithub CopilotGithub Enterprise Server (Ghes)GoHashicorp VaultHcp TerraformKubernetesKubernetes AuthLambdaOrg-Scoped RunnersPkiPythonS3SemgrepSentinelTerraformTerraform Module RegistryTransit Encryption
Artificial Intelligence • Fintech • Greentech • Sales • Software • Travel • Hospitality
Design, build, and maintain scalable AWS infrastructure and developer tooling. Enable teams to own systems, support CI/CD, observability, containers, on-call/incident management, and collaborate with SRE/security to ensure resilient, secure services.
Top Skills:
AWSBashCircleCICloudwatchDatadogDockerEc2EcsGithub ActionsGoIamPulumiPythonRdsS3TerraformVpc
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead and grow an endpoint engineering team to modernize and operate enterprise endpoint platforms for 12,500+ devices. Drive automation, application deployment, device lifecycle management, security by design, vendor/licensing management, and cross-functional partnerships to improve reliability, scalability, and employee experience.
Top Skills:
APIsByodJamf PromacOSMicrosoft IntuneMobile Device ManagementScriptingWindowsZero Trust
Automotive • Hardware • Robotics • Software • Transportation • Manufacturing
Design, implement, and maintain scalable, secure cloud infrastructure with emphasis on AWS, Kubernetes, and Terraform. Build and operate CI/CD pipelines, automate provisioning/configuration, monitor performance, troubleshoot incidents, participate in on-call rotations, create runbooks and documentation, and collaborate cross-functionally to meet application and business requirements.
Top Skills:
AnsibleArtifact RepositoriesAWSAws Secrets ManagerAzureBashChefCi/CdConfluenceDhcpDnsDockerElkGCPGitGrafanaHashicorp VaultHelmIamJenkinsJIRAKubernetesLdapPrometheusPuppetPythonSecurity GroupsTerraformVpcVpn
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
All Filters
Total selected ()
No Results
No Results





















