Wells Fargo Logo

Wells Fargo

Lead Systems Operations Engineer

Posted 5 Days Ago
Be an Early Applicant
Hybrid
Irving, TX
Senior level
Hybrid
Irving, TX
Senior level
Leads site reliability and systems operations for critical consumer-facing platforms. Establishes SLOs, SLIs, and error budgets; improves resilience, observability, automation, monitoring, and recovery. Leads major incident response, root-cause analysis, operational readiness, vendor dependency management, and reliability reporting. Provides technical mentorship and influences architecture, platform engineering, CI/CD, and Infrastructure-as-Code practices across cross-functional teams.
The summary above was generated by AI
About this role:
Wells Fargo is seeking a Lead Site Reliability Engineer (SRE) / Lead Systems Operations Engineer within the Consumer Technology (CT) organization. This role will provide technical leadership for operational excellence, platform reliability, resiliency, observability, and support readiness across critical consumer-facing applications and platforms.
The Lead SRE will serve as a senior technical leader responsible for driving reliability engineering practices, reducing operational risk, improving service availability, and enabling scalable platform operations. This role will partner closely with Application Development, Platform Engineering, Infrastructure teams, Shared Services, and External Vendors to ensure highly resilient, supportable, and observable solutions.
The ideal candidate combines deep technical expertise with strong operational leadership and will play a critical role in advancing Site Reliability Engineering practices across the organization.
In this role, you will support:
Reliability Engineering & Platform Stability
  • Lead reliability initiatives across critical business platforms and customer journeys.
  • Establish and drive Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budget practices.
  • Improve platform resilience through automation, self-healing capabilities, capacity planning, and fault-tolerant designs.
  • Identify and eliminate single points of failure across applications, infrastructure, vendor integrations, and customer flows.
  • Champion engineering solutions that improve availability, scalability, recoverability, and operational maturity.

Incident Management & Operational Excellence
  • Serve as a technical lead during major production incidents, providing coordination, technical direction, and recovery leadership.
  • Drive improvements in Mean Time to Detect (MTTD), Mean Time to Diagnose (MTTDiag), and Mean Time to Recover (MTTR).
  • Lead Root Cause Analysis (RCA) efforts and ensure corrective actions are implemented and tracked to completion.
  • Identify recurring operational patterns and develop preventive solutions to reduce production incidents.
  • Develop and maintain incident playbooks, recovery procedures, and operational readiness standards.

Observability & Monitoring
  • Lead enterprise observability initiatives leveraging Splunk, Grafana, GCP Monitoring, AppDynamics, and related platforms.
  • Define monitoring standards, alerting strategies, dashboards, and customer journey observability solutions.
  • Partner with application and infrastructure teams to improve telemetry, logging, tracing, and synthetic monitoring capabilities.
  • Develop actionable operational metrics and executive-level reliability reporting.

Automation & Engineering Excellence
  • Drive automation strategies that reduce manual effort and improve operational consistency.
  • Design and implement self-service operational capabilities and automated recovery solutions.
  • Utilize AI-assisted tools and engineering practices to improve incident detection, diagnosis, and remediation workflows.
  • Promote Infrastructure-as-Code (IaC), CI/CD best practices, and platform engineering principles.

Vendor & Dependency Management
  • Partner with internal and external service providers to improve reliability, support responsiveness, and recovery performance.
  • Evaluate vendor operational performance and contribute to service improvement initiatives.
  • Establish and monitor operational readiness expectations for critical vendor dependencies.
  • Drive resilience planning and support strategies for third-party integrations.

Technical Leadership
  • Provide technical leadership and mentorship to SREs, Systems Operations Engineers, and Platform Support Engineers.
  • Lead technical reviews, operational readiness assessments, and production support governance activities.
  • Influence architecture decisions to ensure supportability, resiliency, observability, and operational sustainability.
  • Collaborate with engineering leaders to establish and mature Site Reliability Engineering practices across Consumer Technology.
Required Qualifications:
  • 5+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
  • 5+ years of Site Reliability Engineering, Platform Engineering, Production Support, or equivalent experience demonstrated through work experience, military experience, training, or education.
  • 5+ years supporting mission-critical production applications in large enterprise environments.
  • 3+ years leading major incident management, operational support, or reliability engineering initiatives.
  • 3+ years of experience with observability and monitoring platforms such as Splunk, Grafana, AppDynamics, Dynatrace, GCP Monitoring, or similar technologies.
  • 2+ years of experience driving automation, operational improvements, and reliability initiatives.
  • 3+ years of experience supporting distributed systems, cloud-based platforms, infrastructure, networking, and application architectures.
  • 1+ year of experience supporting highly regulated or customer-facing financial services platforms
Desired Qualifications:
  • Consumer Technology, Credit Card, Payments, Lending, or Digital Banking experience.
  • Experience implementing Site Reliability Engineering (SRE) principles, SLOs, SLIs, and Error Budgets.
  • Experience with DevOps, CI/CD, Infrastructure-as-Code, and cloud-native architectures.
  • Experience with AI-assisted engineering, incident management automation, or observability platforms.
  • Strong executive communication and stakeholder management skills.
  • Experience leading cross-functional technical teams without direct authority.
  • Experience supporting vendor governance and third-party operational readiness initiatives.
Job Expectations:
  • Relocation assistance is not provided for this position
  • Visa sponsorship is not available for this position
  • Position requires onsite presence at one of the posted Wells Fargo locations.
Locations:
  • 401 W. Las Collinas Blvd, Irving, Texas
  • 300 S. Brevard St. Charlotte, North Carolina
  • 2600 S. Price Rd. Chandler, Arizona
Posting End Date:
10 Sep 2026
*Job posting may come down early due to volume of applicants.
We Value Equal Opportunity
Wells Fargo is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other legally protected characteristic.
Employees support our focus on building strong customer relationships balanced with a strong risk mitigating and compliance-driven culture which firmly establishes those disciplines as critical to the success of our customers and company. They are accountable for execution of all applicable risk programs (Credit, Market, Financial Crimes, Operational, Regulatory Compliance), which includes effectively following and adhering to applicable Wells Fargo policies and procedures, appropriately fulfilling risk and compliance obligations, timely and effective escalation and remediation of issues, and making sound risk decisions. There is emphasis on proactive monitoring, governance, risk identification and escalation, as well as making sound risk decisions commensurate with the business unit's risk appetite and all risk and compliance program requirements.
Candidates applying to job openings posted in Canada: Applications for employment are encouraged from all qualified candidates, including women, persons with disabilities, aboriginal peoples and visible minorities. Accommodation for applicants with disabilities is available upon request in connection with the recruitment process.
Applicants with Disabilities
To request a medical accommodation during the application or interview process, visit Disability Inclusion at Wells Fargo .
Drug and Alcohol Policy
Wells Fargo maintains a drug free workplace. Please see our Drug and Alcohol Policy to learn more.
Wells Fargo Recruitment and Hiring Requirements:
a. Third-Party recordings are prohibited unless authorized by Wells Fargo.
b. Wells Fargo requires you to directly represent your own experiences during the recruiting and hiring process.

Similar Jobs at Wells Fargo

Senior level
Fintech • Financial Services
Designs and accelerates AI-enabled capabilities across ITSM offerings, initially Incident, Change, and Problem Management. Translates workflows into scalable solution designs, evaluates deterministic automation, predictive, generative, agentic AI, and API approaches, and conducts prototypes and experiments. Produces architecture and implementation guidance, coaches product and engineering teams, influences senior stakeholders, and ensures solutions meet enterprise standards for scalability, supportability, risk, and measurable business outcomes.
Top Skills: Agentic AiAPIsArtificial IntelligenceGenerative AiItsmMachine LearningServicenow
14 Hours Ago
Hybrid
139K-239K Annually
Senior level
139K-239K Annually
Senior level
Fintech • Financial Services
Lead strategy and business performance for personal loans and line products through portfolio analysis, forecasting, lending economics, pricing, financial evaluation, and strategic planning. Develop business cases and recommendations, guide product design and implementation, define product vision and goals, and influence investment and portfolio decisions. Collaborate with marketing, finance, credit risk, legal, product development, technology, vendors, and executive leadership in a complex matrixed environment.
Top Skills: PythonSASSQL
14 Hours Ago
Hybrid
139K-239K Annually
Senior level
139K-239K Annually
Senior level
Fintech • Financial Services
Leads product strategy, roadmap, and delivery for agentic AI automation across home lending originations. Owns backlogs, prioritization, workflow design, responsible AI guardrails, governance, risk controls, adoption, and performance measurement. Partners with engineering, architecture, design, compliance, operations, vendors, and senior stakeholders to transform complex lending processes and deliver scalable autonomous workflows. Drives change management, agile execution, consensus, and long-term product strategy in a regulated financial-services environment.
Top Skills: Agentic AiAgileAi GovernanceAutonomous AgentsData GovernanceHuman-In-The-Loop DecisioningWorkflow Orchestration

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account