Orion Innovation Logo

Orion Innovation

Senior DevOps Engineer - Observability

Posted 24 Days Ago
Remote
2 Locations
Senior level
Remote
2 Locations
Senior level
Design and implement Datadog-based monitoring solutions, automate observability, and lead integrations with cloud platforms and teams, optimizing performance across enterprise environments.
The summary above was generated by AI

Orion Innovation is a premier, award-winning, global business and technology services firm.  Orion delivers game-changing business transformation and product development rooted in digital strategy, experience design, and engineering, with a unique combination of agility, scale, and maturity.  We work with a wide range of clients across many industries including financial services, professional services, telecommunications and media, consumer products, automotive, industrial automation, professional sports and entertainment, life sciences, ecommerce, and education.

Senior DevOps Engineer - Observability

We are seeking a Senior DevOps Engineer focused on Observability to own and drive the observability strategy in our AWS Java-centric cloud environment. You will be responsible for designing and managing a code-driven Datadog observability platform that ensures full visibility into Java applications, Kubernetes workloads and AWS containerized infrastructure all while optimizing Datadog costs and eliminating unnecessary overhead.

This role requires a deep understanding of Datadog, Java logging and tracing, AWS observability best practices and cost efficiency techniques. The ideal candidate will work closely with SRE, DevOps and Software Engineers to standardize monitoring, alerting, metrics and tracing while ensuring cost-effective observability solutions.

As a Senior DevOps Engineer, you will set observability standards, lead automation efforts and mentor engineers ensuring all monitoring and Datadog configuration changes are implemented Infrastructure-as-Code (IaC). You will also drive cross-functional collaboration, working across engineering, infrastructure and product teams to deliver scalable and cost-effective observability outcomes.

Key Responsibilities

Observability Architecture & Strategy

  • Own and define observability standards for Java applications, infrastructure and Kubernetes workloads
  • Design and maintain Datadog observability stacks using Terraform - all configurations must be code-driven
  • Define and enforce JSON structured logging, distributed tracing and metric collection best practices for Java applications
  • Ensure all microservices have proper observability configurations for logs, traces and metrics
  • Implement service-level objectives (SLOs), SLIs, and SLAs to measure application and system reliability
  • Automate observability testing as part of CI/CD pipelines, ensuring new deployments include proper monitoring and logging

Java Application Logging, APM & Distributed Tracing

  • Enforce log filtering and retention policies to prevent excessive log ingestion and reduce Datadog costs
  • Work directly with Java developers to ensure proper logging, tracing and metrics instrumentation
  • Integrate OpenTelemetry for Java distributed tracing, capturing request flow across microservices
  • Standardize logging practices across services using Logback, Log4j or SLF4J ensuring:
    • Structured JSON logs for easy parsing
    • Correlation IDs for linking logs to traces
    • Error reporting integrates with Datadog alerts
  • Implement Datadog APM (Application Performance Monitoring) for Java-based services ensuring deep JVM observability

Datadog Cost Management & Optimization

  • Own Datadog cost governance, ensuring observability costs remain within budget
  • Monitor Datadog ingestion volumes (logs, traces, and metrics) to prevent overages
  • Implement cost-efficient log filtering, retention and sampling policies to reduce unnecessary logging costs
  • Automate Datadog usage reporting and integrate cost tracking into SRE Observability dashboards
  • Establish automated alerts for unexpected increases in Datadog usage costs and cost contributors
  • Advocate for efficient observability by reducing noisy alerts, redundant logs and low-value traces

Incident Response & Reliability Engineering

  • Lead incident response efforts, ensuring Datadog alerts are actionable and tuned to minimize noise
  • Establish RCA and use Datadog data to identify root causes of application failures
  • Work with application teams to optimize their response protocol based on Datadog insights
  • Implement Datadog Watchdog (AI anomaly detection) for proactive failure prediction

Kubernetes & AWS Observability

  • Deploy Datadog monitoring and logging for AWS-native services: EKS, EC2, Lambda, API Gateway, RDS, DynamoDB, SQS, SNS, Step Functions, VPC Flow Logs
  • Optimize Datadog Kubernetes monitoring, ensuring pod-level tracing and resource utilization tracking
  • Kubernetes events and autoscaling alerts are configured in Datadog

Leadership & Mentorship

  • Proactively drive observability initiatives across teams to ensure alignment, adoption and execution of observability goals
  • Act as technical authority on observability, mentoring engineers on Datadog best practices
  • Work closely with Java developers, helping them implement efficient logging, tracing and monitoring
  • Drive observability as a key SRE function, ensuring every new service is fully instrumented from day one
  • Develop internal training programs on Datadog, Terraform-based observability and AWS monitoring

Required Qualifications

  • 5+ years of experience in DevOps, SRE, observability, monitoring, development or software engineering roles
  • Proficient in Terraform for Infrastructure-as-Code (IaC) – Datadog must be managed as code
  • Solid in Datadog including APM, Logs, Metrics, Tracing and Security Monitoring
  • Coding, scripting and automation skills in Java, Python, Node, Bash or Go
  • Experience integrating observability into CI/CD pipelines (GitLab CI, AWS CodePipeline, GitHub)
  • Extensive experience with AWS services and their observability patterns
  • Monitoring experience (ELK, Prometheus, Grafana, OpenTelemetry, New Relic, Dynatrace, Sysdig)
  • Java development background with experience in:
    • Spring Boot, Java microservices architecture
    • Java logging frameworks (Logback, Log4j, SLF4J)
    • Java APM instrumentation with OpenTelemetry or Datadog APM
  • Proficient understanding of JVM performance tuning, GC monitoring, and thread profiling
  • Experience implementing distributed tracing for Java applications
  • Incident response and on-call experience with a proven track record of reducing MTTR


Orion is an equal opportunity employer, and all qualified applicants will receive consideration for employment without regard to race, color, creed, religion, sex, sexual orientation, gender identity or expression, pregnancy, age, national origin, citizenship status, disability status, genetic information, protected veteran status, or any other characteristic protected by law.

Candidate Privacy Policy

Orion Systems Integrators, LLC and its subsidiaries and its affiliates (collectively, “Orion,” “we” or “us”) are committed to protecting your privacy. This Candidate Privacy Policy (orioninc.com) (“Notice”) explains:

  • What information we collect during our application and recruitment process and why we collect it;
  • How we handle that information; and
  • How to access and update that information.

Your use of Orion services is governed by any applicable terms in this notice and our general Privacy Policy.


Top Skills

Ansible
AWS
Azure
Bash
Datadog
GCP
Kubernetes
Python
Terraform

Similar Jobs

9 Days Ago
Remote
Hybrid
Austin, TX, USA
50K-120K
Senior level
50K-120K
Senior level
Artificial Intelligence • Cloud • Information Technology • Sales • Security • Software • Cybersecurity
Develop and maintain automation for observability and reliability, support Kubernetes infrastructure, and mentor engineers in AWS and operational practices.
Top Skills: AnsibleArgoAWSBashDatadogElasticsearchGitGrafanaGroovyIstioJavaJenkinsKubernetesPodmanPrometheusPythonRubySplunkTerraform
An Hour Ago
Easy Apply
Remote
2 Locations
Easy Apply
186K-258K Annually
Senior level
186K-258K Annually
Senior level
Artificial Intelligence • Fintech • Machine Learning • Social Impact • Software
As a Principal Software Engineer at Upstart, you will lead the Identity Platform team, defining and executing on technical direction, improving security measures, guiding architecture, and mentoring developers. You will design large-scale systems that ensure secure and seamless user experiences while collaborating closely with product and security teams.
An Hour Ago
Remote
USA
147K-174K Annually
Junior
147K-174K Annually
Junior
Cloud • Fintech • Cryptocurrency • NFT • Web3
As a Backend Software Engineer, you will decompose a monolithic Rails app into microservices, scale backend systems, and write high-quality code to enhance the Coinbase retail app experience.
Top Skills: DockerGoMongoDBPostgresRuby on RailsRedshiftRubySinatra

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account