crewAI Logo

crewAI

Software Engineer, Infrastructure & Reliability

Reposted 2 Days Ago
Remote
Hiring Remotely in United States
Mid level
Remote
Hiring Remotely in United States
Mid level
Build and operate CrewAI's infrastructure across AWS, Azure, and GCP. Own containerized deployments, CI/CD pipelines, observability, secrets, networking, databases, and on-call reliability. Harden security, automate deployments and installs, partner with runtime and product engineers, and reduce operational toil through tooling and automation.
The summary above was generated by AI
About CrewAI

CrewAI is the leading framework and enterprise platform for building and orchestrating multi-agent AI systems, powering 300M+ agent executions per month across thousands of companies. The Agent Management Platform is our control plane for deploying, monitoring, governing, and scaling agents in production. This role owns the infrastructure foundation that keeps it reliable, secure, and fast.

The Role

You'll build and operate the platform infrastructure behind CrewAI's cloud and enterprise deployments. You'll work across multiple hyperscalers - AWS, Azure, and GCP. You’ll work on containers, CI/CD, deployment automation, observability, secrets, networking, and runtime reliability. Your job is to make the product and runtime teams faster while making customer’s production environments safer.

This is not a pure DevOps support role. You'll write code, improve systems, design deployment paths, harden production, and build the internal platform that lets CrewAI scale and scale our customer deployments.

What You'll Do
  • Own and improve the infrastructure that runs CrewAI's platform: AWS, ECS/ECR, Docker, Kubernetes/Helm, networking, secrets, databases, Redis, and related services.
  • Build and maintain CI/CD pipelines for build, test, image publishing, migrations, environment promotion, rollbacks, and deploy safety.
  • Improve reliability across cloud and enterprise deployments: health checks, alerting, incident response, capacity planning, recovery paths, and operational runbooks - and own the front-line on-call rotation and its SLAs.
  • Partner with runtime engineers on Celery/FastAPI/Redis workloads and with product engineers on Rails/Solid Queue/Postgres production behavior.
  • Manage production observability and telemetry infrastructure: logs, metrics, traces, dashboards, Sentry/OpenTelemetry plumbing, actionable alerts, and telemetry export to customers' own monitoring systems.
  • Harden security and compliance posture across IAM, workload identity, secrets management, vulnerability scanning, dependency/image hygiene, and least-privilege access.
  • Build the tooling and automation that lets field engineers and customers run self-hosted installs themselves - Helm charts, environment config, release artifacts, pre-flight checks, and install runbooks - so engineering does fewer hands-on installs over time.
  • Reduce operational toil by automating recurring workflows and making deployments boring.

RequirementsWhat We're Looking For
  • Strong infrastructure/platform engineering experience in production SaaS environments.
  • Deep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services.
  • Experience with ECS and/or Kubernetes; Helm experience is a strong plus.
  • Comfort operating PostgreSQL, Redis, background job systems, queues, and web services in production.
  • Strong debugging instincts across app, infra, network, deploy, and dependency layers.
  • Security-minded approach to IAM, secrets, workload identity, vulnerability management, and production access.
  • Ability to write reliable automation in Python, Ruby, Go, Bash, or similar.
  • Calm, rigorous approach to incidents, rollbacks, migrations, and production change management.
Bonus
  • Experience with AI/agent platforms, workflow runtimes, or high-volume async execution systems.
  • Experience supporting enterprise/self-hosted deployments.
  • Terraform or other IaC experience.
  • SRE background: SLOs, incident review, capacity planning, load testing.
  • Familiarity with Rails, FastAPI, Celery, OpenTelemetry, or multi-service observability.

Similar Jobs

10 Hours Ago
In-Office or Remote
Junior
Junior
Angel or VC Firm • Professional Services • Consulting • Financial Services
Join a boutique investment banking team to build and own complex financial models, prepare client-ready presentations, drive M&A and capital-raising processes, manage diligence and coordination with advisors, and interact directly with senior clients. Requires strong analytical, quantitative, and client communication skills and elite execution on lean deal teams.
Top Skills: Claude PluginExcelGaapIndex/MatchMacrosPivot TablesPower QueryPowerPointVlookupWordXlookup
10 Hours Ago
In-Office or Remote
105K-165K Annually
Senior level
105K-165K Annually
Senior level
Manufacturing
Lead design, implement, and manage hybrid on-prem and Azure infrastructure. Oversee storage/backup, virtualization, networking, automation (IaC/scripting), monitoring, security/compliance, cost optimization, and lead/mentor infrastructure engineers while coordinating cross-functional teams and supporting data center operations.
Top Skills: AnsibleAristaAryaka Sd-WanAWSAzureAzure Virtual DesktopAzure VmsBashCisco Router/SwitchCommvaultEmcFortinetGCPHyper-VJuniper Mist WirelessNetappPowershellPythonServicenowTerraformVmware Vsphere
11 Hours Ago
Remote or Hybrid
United States
Entry level
Entry level
Fintech • Machine Learning • Software • Financial Services
This position is an opportunity to fill a form for future job matching after the ASPLOS Conference with IMC Trading.

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account