As a Lead Site Reliability Engineer at JPMorgan Chase within the Asset and Wealth Management, Tech Production and Infrastructure Delivery team, you will be responsible for improving reliability, resilience, and operational performance across a hybrid technology environment spanning modern distributed platforms and mainframe systems.
Job Responsibilities
- Lead adoption and operationalization of SRE practices, including SLIs/SLOs, error budgets, reliability reviews, and blameless post-incident processes.
- Design, implement, and continuously improve monitoring and observability capabilities across metrics, logs, traces, and event telemetry to support faster detection and diagnosis.
- Establish actionable alerting standards, dashboards, and runbooks to improve operational readiness and reduce noise.
- Drive automation initiatives (self-service, self-healing, automated remediation, CI/CD operational controls, and standardized tooling) to reduce manual effort and improve consistency.
- Identify, measure, and reduce operational toil through process optimization, tooling enhancements, and platform improvements.
- Improve incident management practices, including incident response coordination, escalation paths, and continuous improvement based on root cause analysis.
- Support capacity planning, performance engineering, and resilience testing to strengthen availability and service stability.
- Partner with application, infrastructure, and operations teams across distributed and mainframe domains to standardize reliability patterns and operational controls.
- Contribute to governance and operational excellence, including documentation, control evidence where applicable, and operational health reporting.
Required qualifications, capabilities and skills
- Relevant experience in Site Reliability Engineering, Production Engineering, Infrastructure Engineering, or a similar reliability-focused role, including leadership of technical initiatives.
- Strong knowledge of operating and supporting distributed systems in production (e.g., Linux, networking, middleware, containers and/or cloud platforms).
- Hands-on experience with monitoring/observability platforms and practices (metrics, logs, traces), including dashboarding and alert engineering.
- Demonstrated ability to automate operational workflows using one or more scripting/programming languages (e.g., Python, Go, Shell) and standard automation approaches (CI/CD, infrastructure-as-code).
- Experience supporting or integrating mainframe systems into enterprise operations (monitoring, incident response, operational processes).
- Strong communication and stakeholder management skills, with ability to lead cross-team reliability improvements.
Preferred Qualifications
- Experience establishing SLIs/SLOs and reliability reporting at service or platform level.
- Familiarity with ITSM/incident tooling, on-call operations, and operational maturity improvements.
- Experience with resilience patterns (graceful degradation, failover, rate limiting) and reliability testing (chaos testing, load/performance testing).
- Exposure to regulated or high-control environments and operational risk management practices.
We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.
We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.
JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/Veterans
Similar Jobs
What you need to know about the Boston Tech Scene
Key Facts About Boston Tech
- Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
- Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
- Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
- Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

