Jackpocket Logo

Jackpocket

Lead Site Reliability Engineer

Posted 4 Hours Ago
Be an Early Applicant
Hybrid
Boston, MA, USA
148K-185K Annually
Senior level
Hybrid
Boston, MA, USA
148K-185K Annually
Senior level
Leads reliability practices across Infrastructure Engineering by defining Service Level Objectives, Service Level Indicators, error budgets, observability standards, and reporting. Partners with engineering teams to connect infrastructure reliability to application and platform outcomes, analyzes distributed-system failure modes, builds Datadog dashboards and monitoring, and influences reliability strategy through technical guidance, mentoring, and executive communication.
The summary above was generated by AI

At DraftKings, AI is becoming an integral part of both our present and future, powering how work gets done today, guiding smarter decisions, and sparking bold ideas. It’s transforming how we enhance customer experiences, streamline operations, and unlock new possibilities. Our teams are energized by innovation and readily embrace emerging technology. We’re not waiting for the future to arrive. We’re shaping it, one bold step at a time. To those who see AI as a driver of progress, come build the future together.

The Crown Is Yours

As a Lead Site Reliability Engineer, you’ll set the reliability standard across our Infrastructure Engineering organization. You’ll define how we measure reliability for critical services, partnering with engineering teams to build and refine Service Level Objectives that connect infrastructure performance to the experiences our platforms support. You’ll turn complex telemetry into clear, actionable insights that help teams and senior leaders make better decisions about reliability, risk, and priorities. As an individual contributor, you’ll lead through technical expertise and influence, shaping a consistent reliability practice across the organization.

What you’ll do as a Lead Site Reliability Engineer
  • Lead and mature the Service Level Objective development process across Infrastructure Engineering, establishing clear frameworks and standards for setting meaningful reliability targets.

  • Partner with engineering teams to design and implement Service Level Objectives, beginning with the critical user journeys each service supports and translating them into measurable indicators, targets, and error budget policies.

  • Review and refine existing reliability objectives to keep them aligned with changing customer impact, technical dependencies, and business priorities.

  • Connect infrastructure reliability targets to the application and platform experiences they support, making dependencies and their impact on end-user experience clear and measurable.

  • Build reporting processes and tooling that provide a clear view of reliability across critical components, translating technical signals into actionable insights for Senior Managers and Directors.

  • Influence reliability practices across teams through technical guidance, design reviews, mentoring, and collaboration with engineering partners.

  • Help teams distinguish meaningful service degradation from true downtime by applying thoughtful measurement strategies to complex distributed systems.

What you’ll bring
  • A Bachelor’s Degree in Computer Science or a related field, or equivalent relevant education, experience, and training.

  • At least 7 years of experience in Site Reliability Engineering, including hands-on experience defining and operationalizing Service Level Objectives, Service Level Indicators, and error budgets at scale.

  • Deep experience with observability platforms such as Datadog, including building dashboards, monitors, and reporting from metrics and logging pipelines.

  • Experience connecting infrastructure-level reliability objectives to application or platform-level outcomes and evaluating how technical dependencies affect end-user experience.

  • Strong knowledge of distributed systems and the failure modes that can make reliability measurement complex, with the ability to assess what technical signals truly represent.

  • Proven ability to influence across engineering teams, translate reliability concepts for technical and non-technical audiences, and drive alignment without direct authority.

  • Excellent written and verbal communication skills, including experience developing and presenting reliability reporting to Senior Managers, Directors, and cross-functional stakeholders.

  • Working knowledge of cloud and infrastructure environments such as Amazon Web Services, Kubernetes, and on-premise systems, with the technical depth to partner effectively with the teams operating them.

Join Our Team

We’re a publicly traded (NASDAQ: DKNG) technology company headquartered in Boston. As a regulated gaming company, you may be required to obtain a gaming license issued by the appropriate state agency as a condition of employment. Don’t worry, we’ll guide you through the process if this is relevant to your role.

The US base salary range for this full-time position is 148,000.00 USD - 185,000.00 USD, plus bonus, equity, and benefits as applicable. Our ranges are determined by role, level, and location. The compensation information displayed on each job posting reflects the range for new hire pay rates for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific pay range and how that was determined during the hiring process. It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.

Similar Jobs

3 Days Ago
Remote or Hybrid
United States
111K-180K Annually
Senior level
111K-180K Annually
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads the architecture, modernization, optimization, reliability, and performance of enterprise IBM z/OS Db2 systems. Provides technical guidance across development and operations teams, automates DDL/DML processes, supports resilience and business continuity testing, resolves Db2 incidents, analyzes performance telemetry, improves SQL and database design, and partners with architects and stakeholders on enterprise technology strategy and hybrid-cloud modernization.
Top Skills: AnsibleCloud IntegrationDdlDevOpsDmlIbm Db2Ibm Z/OsOpenshiftPythonRed Hat AnsibleRmfSmfSQLZlinux
5 Days Ago
Remote or Hybrid
United States
111K-180K Annually
Senior level
111K-180K Annually
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads architecture, modernization, optimization, and reliability initiatives for mainframe CICS, MQ, and z/OS Connect environments. Provides technical direction across development and operations teams, establishes governance and change processes, tunes performance using telemetry, resolves incidents, and develops modernization roadmaps. Collaborates with stakeholders and enterprise architects to deliver secure, scalable, high-availability solutions while evaluating automation, cloud integration, and AI technologies.
Top Skills: AnsibleCicsCobolDevOpsIbm MqIbm Z/OsOpenshiftPythonRed Hat Ansible Automation PlatformZ/Os ConnectZlinux
5 Days Ago
Remote or Hybrid
United States
111K-180K Annually
Senior level
111K-180K Annually
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads the architecture, modernization, resilience, security, and performance optimization of enterprise mainframe environments. Responsibilities include z/OS performance tuning, WLM and RACF administration, business continuity planning, automation, technical governance, incident resolution, stakeholder collaboration, and guidance of cross-functional engineering and operations teams. The role also evaluates cloud, DevOps, AI, and hybrid IT technologies for mainframe transformation.
Top Skills: AnsibleCsmGlobal MirrorIbm Z/OsMetro MirrorOpenshiftPr/SmPythonRacfRed Hat Ansible For Ibm Z CollectionsRmfSmfWlmZlinux

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account