Maximum of 25 job preferences reached.
Top Hybrid DevOps & Platform Engineering Jobs in Boston, MA
Aerospace • Artificial Intelligence • Computer Vision • Machine Learning • Natural Language Processing • Software • Defense
The Cloud Engineer will develop and maintain cloud infrastructure, oversee deployment of cloud services, and ensure compliance with security guidelines. Candidates should possess strong communication skills and be eager to learn new technologies. Experience with AWS, Azure, GCP, Terraform, and Agile development is required.
Top Skills:
AWSAzureCloud FormationContainer Orchestration ServicesGCPMicroservice ArchitectureTerraform
Artificial Intelligence • Cloud • Insurance • Software • Database • Conversational AI • Generative AI
As a Senior Software Engineer II (DevOps), you will design and operate cloud infrastructure, build AI/ML infrastructure, enforce infrastructure-as-code standards, and improve developer experience while coordinating across teams and handling compliance in a regulated environment.
Top Skills:
AirflowAWSBedrockCloudwatchDagsterDatadogDbtDynamoDBEcsGoLambdaPythonRedshiftS3SagemakerTerraformTypescript
Artificial Intelligence • Big Data • Healthtech • Software • Biotech
Design, architect, and operate secure cloud infrastructure (primarily Azure) at scale. Lead Kubernetes implementations and platform reliability, mentor junior engineers, troubleshoot and perform root-cause analysis, and drive cloud infrastructure projects from requirements through technical execution.
Top Skills:
AnsibleAzureBashComputeElkGrafanaKubernetesLinuxLokiNetworkingPrometheusPythonShellStorageTerraformVirtualization
Reposted 21 Hours AgoSaved
Easy Apply
Easy Apply
Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Lead engineering for Cloud Observability across multiple cloud providers, managing ~40 engineers via Engineering Managers. Shape roadmap with product leadership, drive AI adoption, navigate cross-team dependencies, build and retain talent in NYC/Boston/Paris, mentor EMs, and participate in on-call rotations.
Top Skills:
AIAWSAzureDatadogEvent-Driven SystemsGCPGpu Cloud ProvidersIngestion PipelinesOciTelemetryTime-Series
Consumer Web • eCommerce • Software
Design and deliver developer tooling, core frameworks, and templates to improve engineering productivity. Advance agentic AI tooling (Claude Code, plugins, MCP), strengthen authentication/identity (Keycloak, OAuth2/OIDC), and build/operate services on Java, Spring Boot, AWS, and Kubernetes. Define standards, improve reliability, and partner across teams to drive adoption.
Top Skills:
AWSClaude CodeJavaJavaScriptKeycloakKubernetesMcpNode.jsOauth2OidcReactRemixSpring Boot
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead the design, automation, and scaling of global compute infrastructure across data centers, cloud, and on-prem. Operate GitOps with Rancher Fleet/Flux/Helm, build self-healing tooling, own cluster autoscaling and capacity strategy, define SLOs using Datadog, and participate in on-call rotation while mentoring peers.
Top Skills:
AWSContainerdDatadogDockerFluxGCPGitopsGoHelmHpaInfrastructure As Code (Iac)KarpenterKedaKubernetesLinuxNutanixPythonRancher FleetVsphere
Reposted 2 Days AgoSaved
Fintech • Machine Learning • Payments • Software • Financial Services
Lead design, development, deployment, and support of foundational AI systems including foundation model training, LLM inference, similarity search, guardrails, evaluation, and observability. Optimize large-scale AI performance (cost, latency, throughput), partner cross-functionally, and contribute to technical vision and roadmap.
Top Skills:
Aws UltraclustersC#C++GoHuggingfaceJavaLlm InferenceNemo GuardrailsPythonPyTorchScalaSimilarity SearchVectordbs
Reposted 3 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Cloud • Information Technology • Security • Software • Cybersecurity
As a Staff Site Reliability Engineer, you'll oversee Zscaler production data center services, optimize code, and ensure cloud service availability and performance. Collaborate with cross-functional teams to improve processes and resolve escalated issues.
Top Skills:
BashDnsFirewallsGrafanaHTTPIcmpLoad BalancingNagiosOsi ModelPrometheusPythonTcp/Ip
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Design, implement, and operate cloud-native HPC and ML/AI infrastructure (AWS/GCP). Lead containerization, IaC (Terraform/CloudFormation), orchestration (Kubernetes/EKS/GKE), monitoring, and migration. Collaborate with HPC specialists to automate deployments, optimize performance, and maintain platform reliability, security, and cost efficiency.
Top Skills:
ApptainerAWSAws ParallelclusterCloudFormationCloudwatchDockerEksGCPGkeGoogle Cloud Cluster ToolkitGrafanaKubernetesLinuxNvidia Gpu ComputingOpen On DemandPrometheusSingularitySlurmTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top hybrid Companies in Boston, MA Hiring Engineering Roles
See AllAll Filters
Total selected ()
No Results
No Results





















