BJ's Wholesale Club Logo

BJ's Wholesale Club

Principal Engineer - Major Incident Response & ITIL Platform

Posted 5 Days Ago
Be an Early Applicant
In-Office
Marlborough, MA
138K-176K Annually
Senior level
In-Office
Marlborough, MA
138K-176K Annually
Senior level
Owns and evolves the enterprise Major Incident Response, Post-Incident Review, and Problem Management programs. The role engineers ServiceNow ITSM workflows, automations, integrations, dashboards, SLA/SLO tracking, and incident escalation processes. It serves as Incident Commander for high-severity events, leads RCA and continuous improvement, manages an offshore operations team, and communicates program health to executive stakeholders. The position requires participation in on-call rotations and hybrid work onsite in Marlborough, Massachusetts.
The summary above was generated by AI

A World-Class Team

BJ’s Wholesale Club is powered by more than 30,000 team members who make a real impact every day. Whether you're stocking shelves, solving problems or shaping strategy, your work helps families save on what matters most.

We’re a team built on purpose and opportunity. Join us and be part of something meaningful.

Why You’ll Love Working at BJ’s

At BJ’s Wholesale Club, our team members are at the heart of everything we do. That’s why we offer a comprehensive benefits package designed to support your health, well-being and future – both on and off the job. When you grow, we grow.

Here’s just some of what you can look forward to:

  • Weekly Pay: Get paid every week so that you can manage your money on your terms.
  • Free BJ’s Memberships: Enjoy a complimentary The Club Card Membership, plus a free Supplemental Membership for someone in your household.*
  • Generous Paid Time Off: Take the time you need with vacation, personal, sick days, holidays, bereavement, and jury duty leave.*
  • Flexible and Affordable Health Benefits: Choose from three medical plans, and access optional dental, vision, Health Savings Account (HSA), and flexible spending account options to fit your lifestyle.*
  • 401(k) Retirement Savings Plan: Build your financial future with a company match (available to team members 18 and older).*
  • Employee Stock Purchase Plan:  Accumulate funds through after-tax payroll deductions that can be used to purchase shares of BJ’s common stock at a 15% discount.*

*Eligibility requirements vary by position.

Position Overview:

The Principal Engineer, Major Incident Response & ITIL Platform Lead is a senior individual contributor and program leader who combines deep technical expertise with operational discipline. This role owns the design, configuration, and continuous evolution of the ITIL practice — including Major Incident Response (MIR), Post-Incident Review (PIR), and Problem Management — while serving as a hands-on engineer within ServiceNow and adjacent tooling platforms.

Unlike a traditional SDM role, this position is explicitly technical: you will architect workflows, build automation, instrument observability, and drive platform maturity across stores, distribution centers, and digital environments. You will also lead a high-performing offshore team and act as the primary program authority during high-severity events — bridging the gap between engineering execution and executive communication.

Key Responsibilities:

Major Incident Response (MIR) Program Leadership

  • Own and operate the MIR program end-to-end — from playbook authorship to real-time bridge command — for incidents impacting stores, DCs, POS, fuel, e-commerce, and membership systems.
  • Serve as Incident Commander during P1/P2 events, driving technical triage, stakeholder communication, and escalation decisions under pressure.
  • Design and maintain a universal MIR playbook with consistent execution standards 24x7, including on-call rotations for nights, weekends, and holidays.
  • Establish leadership notification templates, technical bridge protocols, and business-facing communication cadences during major incidents.
  • Instrument incident severity classification logic, auto-routing, and escalation thresholds directly within ServiceNow.

Post-Incident Review & Postmortem Excellence

  • Own the end-to-end PIR lifecycle — blameless, data-driven reviews completed within SLA — and enforce action-item closure rigor.
  • Build and maintain an enterprise-wide RCA library, problem signatures, and trend intelligence within ServiceNow's CMDB and Problem Management modules.
  • Partner with SRE and Software Engineering to translate RCA findings into reliability-driven design improvements and automated runbooks.
  • Configure and manage PIR workflows, SLA timers, and notification rules natively in ServiceNow — no manual handoffs.

ITIL Platform Engineering & ServiceNow Ownership

  • Act as a hands-on technical owner of ServiceNow ITSM modules: Incident, Problem, Change, and Event Management.
  • Design and build ServiceNow workflows, business rules, UI policies, Flow Designer automations, and integration spokes connecting monitoring platforms (Dynatrace, Splunk, PagerDuty/AlertOps, etc.).
  • Develop and maintain custom dashboards, real-time KPI reporting, and SLA/SLO tracking within ServiceNow Performance Analytics.

Problem Management

  • Own the Problem Management lifecycle: identification, logging, root cause investigation, routing, and verified resolution.
  • Surface recurring incident patterns from trend analysis and feed intelligence back into MIR and Service Excellence programs.
  • Ensure complete, accurate, and timely documentation of all Problems in ServiceNow with appropriate categorization and linkage to incidents and changes.

Continuous Improvement & Team Leadership

  • Lead and develop a high-performing offshore operations team, setting clear goals aligned to MIR and ITIL program objectives.
  • Drive a culture of automation-first thinking: identify manual toil and eliminate it through ServiceNow scripting, Flow Designer, and third-party integrations.
  • Conduct regular retrospectives, process audits, and tooling reviews; translate findings into prioritized improvement backlog items.
  • Present program health, metrics, and roadmap updates to senior IT and business leadership.

Key Outcomes:

  • Faster stabilization of high-severity events through structured, technically informed incident command.
  • Measurable reduction in repeat incidents via high-quality, action-tracked RCAs.
  • A mature, automated ServiceNow platform that minimizes manual effort and accelerates response and reporting.
  • Predictable, trust-building communication to business stakeholders during and after major incidents.
  • Continuous improvement embedded into operational DNA — not a periodic exercise.

KPIs & Success Metrics:

Implement and manage service level agreements (SLAs and SLOs) to meet organizational goals and user expectations.  

  • Mean Time to Acknowledge (MTTA) and Mean Time to Resolve (MTTR) for P1/P2 incidents.
  • Postmortem SLA compliance (e.g., 100% PIR completion within 5 business days).
  • Action item closure rate from PIRs within agreed timelines.
  • Reduction in repeat incidents (measured quarterly).
  • Problem Management throughput (number of problems logged, analyzed, and resolved).
  • Leadership communication SLA adherence during major incidents.
  • Continuous improvement initiatives delivered (e.g., automation, process optimization).
  • Stakeholder satisfaction scores from incident and problem management processes.

Requirements:

Education & Certifications

  • Bachelor's degree in Computer Science, Information Systems, or equivalent experience.
  • ITIL 4 Strategic Leader or Managing Professional certification strongly preferred; ITIL Expert acceptable.
  • ServiceNow certifications (CSA, CIS-ITSM, or CIS-Event Management) highly desirable.

Experience

  • 8+ years in IT Service Management with a strong technical bias — hands-on platform work, not just process governance.
  • 5+ years of direct experience administering or engineering ServiceNow (workflow design, scripting, integrations, Performance Analytics).
  • Proven track record leading Major Incident Response in large-scale retail, e-commerce, or distributed digital environments.
  • Experience owning Problem Management and PIR programs with measurable outcomes (repeat incident reduction, MTTR improvement).
  • Demonstrated ability to manage and develop onshore/offshore teams in a follow-the-sun operations model.

Work Environment:

  • Hybrid working model: 3 days onsite (Tue, Wed, Thu), 2 days remote (Mon, Fri).
  • Occasional travel to company locations or industry events.
  • Flexibility in working hours to accommodate global operations and time zone differences.
  • Participation in Major Incident on-call rotation.

Technical Skills

  • Deep ServiceNow platform expertise understanding platform mechanics that drive configuration and administration.
  • Working knowledge of tools such as Dynatrace, Splunk, AlertOps/PagerDuty, or equivalent; ability to build event-to-incident automation bridges. Observability & Monitoring:
  • MS Teams, Jira — including integration design with ServiceNow. Collaboration & Comms:
  • ServiceNow Performance Analytics, dashboard design, SLA/SLO instrumentation. Data & Reporting:
  • Comfort with scripting (JavaScript, Python, or PowerShell) to accelerate toil elimination. Scripting / Automation:

Leadership & Soft Skills

  • Able to command a major incident bridge — calm, decisive, technically credible under pressure.
  • Executive-level communication: concise, audience-aware, and trustworthy during crises.
  • Program management discipline: roadmaps, metrics, stakeholder alignment, and backlog ownership.
  • Growth mindset with a bias toward automation and measurable improvement.

This is a hybrid role. Tuesday through Thursday are in-office days at BJ's Club Support Center in Marlborough, MA and Monday and Friday are remote days.

In accordance with the Pay Transparency requirements, the following represents a good faith estimate of the compensation range for this position. At BJ’s Wholesale Club, we carefully consider a wide range of non-discriminatory factors when determining salary. Actual salaries will vary depending on factors including but not limited to location, education, experience, and qualifications. The pay range for this position is $137,500.00 - $175,500.00

We recognize the growing role of AI tools, including ChatGPT, and value familiarity with them. That said, we want to hear from your authentic self. Your application should reflect your own skills, experiences, and insights rather than AI-generated responses.

Similar Jobs

An Hour Ago
Remote or Hybrid
179K-322K Annually
Senior level
179K-322K Annually
Senior level
Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Lead the strategic and operational direction of a specialized casualty claims team managing severe-exposure portfolios. Oversee claims managers, complex litigation, coverage strategy, reserves, settlements, legal vendors, and portfolio analytics. Collaborate with legal, underwriting, actuarial, reinsurance, and executive stakeholders to manage emerging risks, litigation trends, loss costs, and claim outcomes. Develop protocols and best practices for environmental, mass tort, toxic tort, latent injury, and other complex claims.
Top Skills: Guidewire Claims System
2 Hours Ago
Remote or Hybrid
USA
100K-155K Annually
Senior level
100K-155K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Build and operate AWS GovCloud data infrastructure across development, preproduction, and production environments. Establish PostgreSQL platforms, infrastructure-as-code with Chef, BI infrastructure, identity integrations, monitoring, observability, CI/CD, and disaster recovery. Support data pipelines, troubleshoot incidents, optimize databases, and ensure FedRAMP and FISMA compliance. Collaborate with security, compliance, DevOps, and data engineering teams while delivering greenfield infrastructure from architecture through production.
Top Skills: AWSAws GovcloudBashChefCi/CdCloudwatchDatadogEltETLGitGitlabGoogle SamlJenkinsNagiosPostgresPythonRubySsoTableau ServerVpc
2 Hours Ago
Remote or Hybrid
USA
140K-215K Annually
Senior level
140K-215K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Design, implement, optimize, and maintain scalable hybrid multi-cloud Kubernetes platforms operating at massive scale. Integrate open-source technologies, improve platform reliability, provide technical direction, mentor engineers, and participate in on-call support. The role requires expertise managing large Kubernetes clusters, observability tools, Linux environments, public clouds, custom data centers, and AI-enabled workflow improvements.
Top Skills: AlertmanagerAWSGCPGoGrafanaHybrid Multi-CloudKubernetesKubernetes OperatorsLinuxOciPrometheusThanos

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account