Federal Reserve System Logo

Federal Reserve System

Lead Site Reliability Engineer

Reposted 2 Days Ago
Be an Early Applicant
In-Office
Boston, MA, USA
147K-234K Annually
Senior level
In-Office
Boston, MA, USA
147K-234K Annually
Senior level
Leads the design and operation of highly available cloud systems, primarily on AWS, using Terraform, GitLab CI/CD, containers, and observability tools. Responsibilities include reliability engineering, incident response, disaster recovery, automation, monitoring, security integration, vulnerability management, and cost optimization. The role builds internal tools, guides architecture, mentors SRE engineers, documents operational processes, and partners with software engineering teams to improve production reliability and compliance.
The summary above was generated by AI
CompanyFederal Reserve Bank of San Francisco

When you join the Federal Reserve—the nation's central bank—you’ll play a key role, collaborating with leading tech professionals to strengthen and protect our economic, financial and payments systems. We invest in contemporary and emerging technology each year to support the Federal Reserve and our economy, and we’re building a dynamic and diverse team for our future.
The Federal Reserve Financial Services portfolio provides technology capabilities that power the U.S. payment systems infrastructure. This portfolio enables critical payment services that are foundational to the nation's financial system, focusing on delivering speed, resilience, and choice to meet evolving marketplace needs.
We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role combines deep technical expertise in software engineering, cloud infrastructure, and DevOps practices to ensure our services meet the highest standards of availability and operational excellence.

Responsibilities

System Reliability & Performance

•    Design, implement, and maintain highly available, scalable, and resilient systems across cloud infrastructure

•    Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance

•    Lead incident response, conduct root cause analysis, and implement preventive measures

•    Develop and maintain disaster recovery and business continuity plans

Infrastructure & Automation

•    Architect and manage cloud infrastructure on AWS using Infrastructure as Code (Terraform)

•    Automate deployment pipelines, monitoring, and operational workflows

•    Optimize cloud resource utilization and cost management

Engineering & Development

•    Build and maintain internal tools and services to improve operational efficiency

•    Collaborate with development teams to implement reliability best practices

•    Conduct code reviews and provide technical guidance on system design

•    Develop monitoring solutions, alerting systems, and observability frameworks

Security & Compliance

•    Integrate security practices into CI/CD pipelines (SAST/DAST)

•    Implement and maintain security controls across infrastructure and applications

•    Ensure compliance with industry standards and regulatory requirements

•    Conduct security assessments and vulnerability management

Leadership & Collaboration

•    Mentor junior SRE team members and promote SRE culture across the organization

•    Partner with software engineering teams to improve system reliability

•    Drive technical initiatives and contribute to architectural decisions

•    Document processes, runbooks, and technical specifications
 

Software Engineering:

  • Strong proficiency in Java, Python, and Node.js
  • Experience with microservices architecture and distributed systems
  • Solid understanding of data structures, algorithms, and design patterns
  • Proficiency in writing clean, maintainable, and testable code

Cloud Infrastructure (AWS):

  • Extensive experience with AWS services including:
    • Compute: Lambda, ECS, EC2, Fargate
    • Storage: S3, EBS, EFS
    • Database: RDS, DynamoDB, Aurora
    • Networking: VPC, Route53, CloudFront, API Gateway
    • Monitoring: CloudWatch, X-Ray
  • AWS certifications (Solutions Architect, DevOps Engineer) preferred

DevOps & CI/CD:

  • Expert-level knowledge of GitLab (CI/CD pipelines, runners, GitOps)
  • Advanced Terraform skills for infrastructure provisioning and management
  • Experience with containerization (Docker) and orchestration (Kubernetes/ECS)
  • Proficiency with configuration management tools

Security:

  • Hands-on experience with SAST (Static Application Security Testing) tools
  • Knowledge of DAST (Dynamic Application Security Testing) methodologies
  • Understanding of security best practices, OWASP Top 10, and compliance frameworks
  • Experience with secrets management and identity access management (IAM)

Monitoring & Observability:

  • Experience with monitoring tools (Grafana, Datadog, New Relic, or similar)
  • Log aggregation and analysis (CloudWatch Logs, Splunk)
  • Distributed tracing with aws X-Ray

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience
  • 7+ years of experience in Site Reliability Engineering, DevOps, or related roles
  • 3+ years in a lead or senior technical position
  • Proven track record of managing large-scale production systems
  • Experience with on-call rotations and incident management
  • GenAI based Applications: Working knowledge of LLMs and agentic applications a plus
  • Experience with serverless architectures and event-driven systems
  • Familiarity with chaos engineering principles and practices
  • Background in Agile/Scrum methodologies
  • Experience with multi-cloud or hybrid cloud environments

The selected candidate will reside within a reasonable commuting distance, as defined by the employing Reserve Bank, and will work full-time onsite.

  • Eligible Locations for Hire:  Boston, MA- New York, NY- Philadelphia, PA- Cleveland, OH- Richmond, VA- Atlanta, GA- Chicago, IL- St. Louis, MO- Minneapolis, MN- Kansas City, MO- Dallas, TX- San Francisco, CA
  • The following Reserve Bank locations are preferred due to the concentration of System IT team members in these locations: San Francisco, and Richmond, VA

Screening:

Due to the nature of access to sensitive information all final offers are subject to the clearance of an enhanced background check. This enhanced screening will require the following items: academic and employment verifications, FBI fingerprint check (criminal and civil cases), credit check, family history, residential records and foreign travel for the previous 7 years, citizenship verification, reference checks, and personal interview with an investigator and can take between 21 – 60 days to clear.


Sponsorship:
Individuals who need immigration sponsorship now or in the future are not eligible for this position.
Must be a U.S Citizen or a Green card holder with intent to become a U.S Citizen.


Base Salary Range: Min: $146,700 Mid: $190,500 Max: $234,300 (Location: San Francisco)

The listed salary is applicable to 12th District/San Francisco. Final offers are determined by factors including the candidate’s qualifications, internal alignment considerations, district assignment, and geographic location.


The Bank is committed to providing reasonable accommodations to individuals with disabilities to participate in the job application or interview process, perform essential job functions and receive other benefits and privileges of employment. The SF Fed is an Equal Opportunity Employer. If you need any assistance or accommodations due to a disability, please let us know at [email protected].

Full Time / Part TimeFull time

Regular / TemporaryRegular

Job Exempt (Yes / No)Yes

Job CategoryInformation Technology Family Group

Work ShiftFirst (United States of America)

The Federal Reserve Banks are committed to equal employment opportunity for employees and job applicants in compliance with applicable law and to an environment where employees are valued for their differences.

Always verify and apply to jobs on Federal Reserve System Careers (https://rb.wd5.myworkdayjobs.com/FRS) or through verified Federal Reserve Bank social media channels.

Privacy Notice

Similar Jobs

2 Days Ago
In-Office
Boston, MA, USA
147K-234K Annually
Senior level
147K-234K Annually
Senior level
Fintech • Information Technology • Payments • Sharing Economy • Financial Services • Cryptocurrency
Leads reliability, scalability, performance, and security for large-scale cloud systems. Designs AWS infrastructure with Terraform, automates CI/CD and operational workflows, establishes SLOs, manages incident response and disaster recovery, develops monitoring and observability solutions, and builds internal tools. Partners with engineering teams on architecture and reliability practices, conducts code reviews, mentors SREs, and supports compliance, vulnerability management, and security integration.
Top Skills: Agentic ApplicationsAmazon Api GatewayAmazon AuroraAmazon CloudfrontAmazon CloudwatchAmazon DynamodbAmazon EbsAmazon Ec2Amazon EcsAmazon EfsAmazon RdsAmazon Route 53Amazon S3Amazon VpcAWSAws FargateAws LambdaAws X-RayChaos EngineeringCi/CdDastDatadogDistributed SystemsDockerEvent-Driven SystemsGitlabGitopsGrafanaIamInfrastructure As CodeJavaKubernetesLlmsMicroservicesNew RelicNode.jsOwasp Top 10PythonSastServerless ArchitecturesSplunkTerraform
6 Days Ago
Hybrid
Senior level
Senior level
Financial Services
Leads architecture and engineering for scalable backup, recovery, immutable-data protection, and recovery-assurance services. Builds AWS cloud-native infrastructure using automation, APIs, infrastructure as code, and CI/CD. Implements SRE practices, observability, security controls, audit evidence, telemetry, and operational readiness. Uses validated AI-assisted workflows, mentors engineers, and coordinates complex cross-functional initiatives.
Top Skills: APIsAWSAws BackupC#Ci/CdCloudFormationCohesityCommvaultDynatraceGithub ActionsGitlabGoGrafanaIamImmutable StorageInfrastructure As CodeJavaJenkinsKubernetesObservabilityOpenshiftPythonSreTerraform
13 Days Ago
Remote or Hybrid
United States
111K-180K Annually
Senior level
111K-180K Annually
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads architecture, modernization, optimization, and reliability initiatives for mainframe CICS, MQ, and z/OS Connect environments. Provides technical direction across development and operations teams, establishes governance and change processes, tunes performance using telemetry, resolves incidents, and develops modernization roadmaps. Collaborates with stakeholders and enterprise architects to deliver secure, scalable, high-availability solutions while evaluating automation, cloud integration, and AI technologies.
Top Skills: AnsibleCicsCobolDevOpsIbm MqIbm Z/OsOpenshiftPythonRed Hat Ansible Automation PlatformZ/Os ConnectZlinux

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account