Berkeley Research Group Logo

Berkeley Research Group

Site Reliability Engineer

Reposted One Month Ago
Remote
Hiring Remotely in USA
130K-160K Annually
Senior level
Remote
Hiring Remotely in USA
130K-160K Annually
Senior level
Design, build, and maintain highly available cloud-native systems. Improve reliability through automation, CI/CD, Kubernetes, observability, and incident management. Collaborate with developers, security, and product teams to define SLOs, implement self-healing, debug production issues, and ensure secure deployments.
The summary above was generated by AI
We do Consulting Differently

Second Sight Solutions, a subsidiary of Berkeley Research Group (BRG), is a health technology company, and our innovative technology reimagines how drug discount data is exchanged, establishing new connections and improving transparency for drug manufacturers and their customers. Our customers and partners trust us to deliver reliable, first-to-market solutions and safeguard the data we receive. We trust our employees, and our culture gives them the freedom to create, collaborate, and grow. Our leaders are industry experts, creative, unafraid to challenge the status quo, and the pioneers of market-changing solutions.

We are seeking a Site Reliability Engineer to design, build, and maintain highly available systems and infrastructure.  The SRE will work closely with software developers and operations teams to improve system reliability, automate processes, and minimize downtime.

Responsibilities

  • Design, implement, and maintain scalable and reliable systems in cloud environments such as Azure Cloud Services.

  • Experience with CI/CD Platforms (GitHub Actions, GitLab CI)

  • Provide operational support for full-stack software applications.

  • Increase system resilience with expert-level coding, bulletproof release, and change management skills.

  • Develop service-level indicators and objectives to automate release validation. 

  • Improve automation and increase the system’s self-healing capability.

  • Collect operating system data and report performance metrics to stakeholders.

  • Ensure security best practices are followed in cloud infrastructure and application deployments.

  • Manage cloud and database system maintenance, debugging production issues as they arise.

  • Improve reliability, quality, and time-to-market of our suite of software solutions.

  • Partner with security and product teams to define and publish policies, processes, and playbooks to facilitate rapid and effective handling of alerts and incidents.

  • Lead incident management processes; respond to outages and service disruptions promptly.

Qualifications:

  • Bachelor’s degree in computer science or similar field.

  • Five years’ experience as a site reliability engineer or similar role.

  • Strong programming skills (Golang, Ruby, Python, or similar)

  • Proven ability to diagnose and monitor performance and reliability issues across the stack.

  • Expertise in Kubernetes.

  • Relevant industry certifications, such as through the Site Reliability Engineering (SRE) Foundation.

  • Proven experience working with cloud-native infrastructure (Azure Cloud Services, AWS, or GCP).

  • Experience working with observability and incident management tools (Datadog, OpsGenie, PagerDuty).

  • Experience scripting operating system tasks with Infrastructure as Code.

  • Impeccable communication skills.

  • Ability to problem-solve in a fast-paced, high-stakes environment.

Candidate must be able to submit verification of his/her legal right to work in the United States, without company sponsorship.

Salary: $130,000 - $160,000
 

About BRG

BRG combines world-leading academic credentials with world-tested business expertise and purpose-built emerging technologies. Our culture centers on agility and connectivity which sets us apart and gets you ahead.  


At BRG, our professionals include specialist consultants, industry experts, renowned academics, and leading-edge data scientists. Together, they bring a diversity of real-world experience, data, and human and artificial intelligence, to economics, disputes, and investigations; corporate finance; and performance improvement services that address the most complex challenges facing organizations across the globe.


Our unique structure nurtures the interdisciplinary relationships that give us the edge, laying the groundwork for more informed insights and more original, incisive thinking.  When paired with our global reach and resources, our diverse perspectives and technical capabilities make us uniquely capable to address our clients’ challenges. We get results because we know how to apply our thinking to your world.


At BRG, we don’t just show you what’s possible. We’re built to help you make it happen. 

BRG is proud to be an Equal Opportunity Employer. Our hiring practices provide equal opportunity for employment without regard to race, religion, color, sex, gender, national origin, age, United States military veteran status, ancestry, sexual orientation, marital status, family structure, medical condition including genetic characteristics or information, veteran status, or mental or physical disability so long as the essential functions of the job can be performed with or without reasonable accommodation, or any other protected category under federal, state, or local law.

Similar Jobs

18 Minutes Ago
In-Office or Remote
United States
85K-193K Annually
Senior level
85K-193K Annually
Senior level
Automotive
Design, build, and operate a global observability platform across hybrid cloud and on-premises environments. Responsibilities include developing monitoring pipelines, infrastructure-as-code, reliability automation, SLI/SLO frameworks, performance optimization, incident response, root-cause analysis, and production troubleshooting. The role partners with engineering teams to improve system resilience, reduce toil, integrate AI/ML for anomaly detection, and establish observability best practices. It also provides technical mentorship and guidance.
Top Skills: Ai/MlAmazon Web ServicesCC++Ci/CdDatadogDockerDynatraceElkGoGoogle Cloud PlatformJ2EeJavaKafkaKubernetesMicroservicesAzureNagiosNew RelicNoSQLOpentofuPrometheusPythonRestful ApisScalaSensuSplunkSpring BootSQLTcp/IpTerraform
19 Minutes Ago
In-Office or Remote
United States
85K-193K Annually
Mid level
85K-193K Annually
Mid level
Automotive
Develop and maintain global monitoring and observability platforms using Go, JavaScript, GCP, Kubernetes, OpenTelemetry, PostgreSQL, and Terraform. Improve reliability, scalability, performance, security, and disaster recovery for cloud services. Responsibilities include troubleshooting production systems, capacity planning, automation, on-call support, incident postmortems, code reviews, documentation, and vulnerability assessments.
Top Skills: Document DatabasesDynatraceGoGoogle Cloud PlatformInfrastructure As CodeJavaScriptKubernetesOpentelemetryPostgresRelational DatabasesTerraform
Yesterday
Remote or Hybrid
OH, USA
Senior level
Senior level
Financial Services
Leads AWS cloud production management and Site Reliability Engineering practices for business-critical banking services. Defines SLOs, SLIs, error budgets, observability, capacity planning, resiliency, and disaster recovery strategies. Directs major incident response, root-cause analysis, and corrective actions; advances infrastructure automation, cloud cost optimization, and multi-region architecture. Establishes production readiness and operational standards while guiding responsible AI-assisted engineering adoption and communicating technical strategy to leadership.
Top Skills: Active-Active SystemsAi-Assisted Software Development ToolsAlertingAnomaly DetectionAutomationAWSCloud ArchitectureCloud InfrastructureDashboardsDisaster RecoveryDistributed TracingInfrastructure As CodeLoggingMetricsMulti-Region ArchitectureObservabilitySre

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account