Kintsugi AI, Inc. Logo

Kintsugi AI, Inc.

Senior Site Reliability Engineer (DevOps)

Reposted 9 Hours Ago
Remote
Hiring Remotely in United States
180K-190K Annually
Senior level
Remote
Hiring Remotely in United States
180K-190K Annually
Senior level
Own and improve the reliability of Kintsugi’s AWS and managed Kubernetes infrastructure. Build automation, internal tools, observability, CI/CD pipelines, and infrastructure-as-code workflows while reducing operational toil. Lead incident response, reliability architecture, developer enablement, cost optimization, security, compliance, and disaster recovery efforts. Work agent-first and partner with Platform Engineering, Product, QA, and software teams to deliver scalable, secure systems.
The summary above was generated by AI
About Kintsugi

Kintsugi is revolutionizing sales tax automation with our AI-powered platform designed specifically for e-commerce and SaaS businesses. Our solution reduces tax preparation time by 75 percent and cuts compliance costs by 50 percent, allowing finance teams to focus on strategic initiatives rather than routine calculations. As we continue to grow and disrupt the tax automation space, we're building a world-class team to help us scale with purpose.

The Role

We're looking for a Senior Site Reliability Engineer (DevOps) to help scale and harden the infrastructure that powers Kintsugi. This role sits at the intersection of software engineering and operations: you'll keep production reliable under real load, and you'll build the tooling and automation that reduce the manual work of doing that, rather than absorbing more of it yourself as we grow.

You'll work closely with Platform Engineering, Product, and QA to design resilient architectures, improve deployment pipelines, and build the internal tools and guardrails that let the team move quickly without sacrificing stability. You'll operate across our managed Kubernetes and AWS-hosted data layer (Postgres, Redis, networking), and you'll be a key force in shaping the reliability and developer-experience foundation of our engineering org.

We're an agentic-coding-first team — coding agents already do real engineering and operations work here, not just autocomplete on the side. We expect this role to build and extend that practice, not just adopt it.

What You'll Do
  • Own the reliability of production infrastructure running on managed Kubernetes and AWS, keeping a high-traffic system up and catching issues before customers do

  • Work agent-first day to day: build, debug, and automate using agentic coding workflows as your default mode, not a fallback tool

  • Build internal tools and automation that eliminate recurring manual work (toil) for the team, rather than just documenting runbooks around it

  • Develop and operate monitoring, alerting, and observability systems (metrics, tracing, logging) across the stack

  • Partner with engineering teams to design for reliability and performance from the start, not bolt it on after incidents

  • Automate infrastructure management through infrastructure-as-code, and improve CI/CD pipelines and local developer workflows

  • Lead and evolve incident response practices, including postmortems and blameless learning

  • Optimize infrastructure for cost efficiency while maintaining high availability and security standards

  • Contribute to security, compliance, and disaster recovery efforts as the platform scales

  • Support developer enablement: build and improve the in-house tooling, local dev workflows, and internal platforms other engineers rely on

What We're Looking For
  • 5-8 years in Site Reliability Engineering, DevOps, or Infrastructure Engineering roles, with real ownership of a production system at meaningful scale — someone who drives reliability and tooling initiatives rather than waiting to be assigned them

  • Fluent working agent-first day to day — directing coding agents to do real engineering work, not just occasional autocomplete

  • Strong foundation in AWS-hosted data and networking services (RDS/Postgres, ElastiCache/Redis, VPC/networking) and experience running workloads on managed Kubernetes

  • A track record of building tools, not just running playbooks: scripts, services, or internal platforms that removed manual work for a team

  • Hands-on experience with CI/CD pipelines and infrastructure-as-code (e.g., Terraform, CloudFormation)

  • Expertise in observability stacks (metrics, tracing, logging) and modern monitoring practices

  • Familiarity with security and compliance in cloud environments (SOC 2, GDPR, etc. a plus)

  • A collaborative mindset with a passion for empowering developers to move fast safely

  • Experience with (or strong interest in) developer enablement — building the internal tooling and platforms that make other engineers more productive

  • Nice to have: experience operating across multiple cloud providers, as our infrastructure footprint expands

Similar Jobs

8 Days Ago
In-Office or Remote
92K-164K Annually
Senior level
92K-164K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Designs and operates secure, reliable Azure cloud platforms using Terraform, GitHub Actions, containers, and automation. Responsibilities include CI/CD, observability, incident response, platform security, vulnerability remediation, disaster recovery, infrastructure troubleshooting, and SRE practices. The role supports production workloads, improves reliability and delivery processes, participates in on-call activities, and mentors engineers while partnering across development, security, architecture, and operations teams.
Top Skills: BashCi/CdCloud SecurityDockerGitGithub ActionsGitopsInfrastructure As CodeKubernetesAzureObservabilityPowershellPythonTerraform
One Month Ago
In-Office or Remote
United States
165K-215K Annually
Senior level
165K-215K Annually
Senior level
Software • Cybersecurity
This role involves managing Kubernetes clusters, cloud infrastructure, and CI/CD pipelines. The engineer will enhance system reliability and efficiency while troubleshooting production issues.
Top Skills: AlertmanagerAWSAzureBashCi/CdDockerElastic StackElasticsearchGCPGoGrafanaHelmKafkaKubernetesLokiMongoDBOciPrometheusPythonRedisSparkTerraform
One Month Ago
Remote
United States
Senior level
Senior level
Big Data
You will manage AWS infrastructure, automate deployments, debug application issues, and improve the operational health of Metabase Cloud.
Top Skills: AWSDatadogGoGrafanaKubernetesPrometheusPythonTerraform

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account