Bright Vision Technologies Logo

Bright Vision Technologies

Observability Engineer

Posted 8 Days Ago
Be an Early Applicant
In-Office
Westford, MA
100K-160K Annually
Expert/Leader
In-Office
Westford, MA
100K-160K Annually
Expert/Leader
Design, operate, and scale enterprise observability platforms (metrics, logs, traces, alerts). Architect Prometheus/Grafana and commercial tools, define instrumentation standards, SLOs/SLIs, alerting, tracing pipelines, cost management, self-service tooling, incident readiness, and mentor teams.
The summary above was generated by AI
Observability Engineer - Remote
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: Observability Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $100,000–$160,000 Annually
Experience Required: 12+ years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Job Summary:
We are looking for an Observability Engineer to design and operate the metrics, logging, tracing, and alerting platforms that give engineering teams confidence in the systems they run. The role spans the full observability stack — from collection agents and pipelines to long-term storage, dashboards, and alerting workflows — with a strong focus on usability, signal quality, and operational ROI. The ideal candidate has built and operated observability platforms at scale, understands the trade-offs between open-source and SaaS approaches, and can translate noisy telemetry into actionable insight for both engineers and business stakeholders.
Key Responsibilities
  • Design and operate enterprise-grade observability platforms covering metrics, logs, traces, events, and synthetic monitoring.
  • Architect Prometheus / Thanos / Mimir, Grafana, Loki, Tempo, OpenTelemetry, and Datadog deployments for high availability and scale.
  • Develop standards for service instrumentation, including OpenTelemetry adoption, metric naming, label cardinality, and structured logging conventions.
  • Define and enforce SLOs, SLIs, and error budgets, and build the dashboards and alerts that operationalize them.
  • Build alerting strategies that minimize noise, surface actionable signals, and integrate cleanly with on-call workflows in PagerDuty, Opsgenie, or similar tools.
  • Operate large-scale time-series and log storage platforms, balancing retention, query performance, and cost.
  • Design distributed tracing pipelines and help teams use traces to diagnose latency and reliability issues.
  • Develop self-service tooling, paved-road libraries, and templates that make adoption of observability standards easy for product teams.
  • Drive cost management and label-cardinality discipline across the observability estate.
  • Lead incident response readiness improvements through better dashboards, alerting hygiene, and post-incident analysis tooling.
  • Partner with SRE and platform teams to integrate observability into deployment pipelines, canary analysis, and progressive delivery workflows.
  • Evaluate and recommend observability vendors and open-source tools based on cost, capability, and operational maturity.
  • Mentor engineering teams on observability fundamentals, debugging techniques, and SLO-driven operations.
  • Maintain documentation, onboarding guides, and runbooks for the observability platform.
Required Qualifications
  • Bachelor’s degree in Computer Science or a related field.
  • Five or more years of experience in SRE, platform engineering, or observability roles.
  • Deep hands-on experience with Prometheus, Grafana, and at least one major commercial observability platform such as Datadog, New Relic, or Splunk.
  • Strong understanding of OpenTelemetry, distributed tracing, and structured logging.
  • Proficiency in at least one general-purpose language such as Go, Python, or Java.
  • Experience operating high-cardinality, high-throughput metrics and log pipelines.
  • Strong understanding of SLOs, error budgets, and SRE principles.
  • Experience integrating observability with CI/CD and incident management tooling.
  • Solid grasp of Linux internals, networking, and container platforms.
  • Excellent communication and collaboration skills.
Preferred Qualifications
  • Experience with Thanos, Mimir, Cortex, Loki, or Tempo at scale.
  • Contributions to OpenTelemetry or observability open-source projects.
  • Familiarity with eBPF-based observability tooling.
  • Experience driving observability cost optimization initiatives.
  • Exposure to regulated environments with audit-grade logging requirements.
How to Apply
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3899. Learn more about Bright Vision Technologies at www.bvteck.com.
Bright Vision Technologies is an Equal Opportunity Employer.
 

Similar Jobs

2 Days Ago
In-Office or Remote
6 Locations
Expert/Leader
Expert/Leader
Software
Serve as the technical lead for Acceldata's Hadoop observability and open data platform in pre-sales: deliver demos, design solutions, run POCs, craft technical proposals/SOWs, collaborate with sales/product/engineering, and support closing deals while educating internal teams and producing technical documentation.
Top Skills: AcceldataAirflowSparkAWSAzureDatabricksGCPHadoopHbaseHdfsHiveJupyterKafkaKerberosRangerSnowflakeSparkTrinoYarnZookeeper
4 Days Ago
In-Office
Boston, MA, USA
230K-270K Annually
Expert/Leader
230K-270K Annually
Expert/Leader
Information Technology • Software • Database
The Principal Software Engineer will lead architectural decisions, mentor engineers, optimize system performance and reliability, and drive technical direction for LangChain's observability and evaluations platform.
Top Skills: AWSAzureClickhouseGCPGoPostgresPythonReactRedisTypescript
18 Days Ago
In-Office or Remote
United States
Senior level
Senior level
Cloud • Information Technology • Software • Infrastructure as a Service (IaaS)
Build ingestion pipelines for logs and metrics, scalable alerting engines, and observability APIs. Interface with product teams and develop microservices using Golang and Rust.
Top Skills: AnsibleGoGraphQLGrpcRustTerraformTypescript

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account