EverOps Logo

EverOps

Senior Infrastructure Engineer

Posted 14 Days Ago
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Lead design, implementation, and troubleshooting of enterprise network and hybrid cloud infrastructure. Build and automate infrastructure-as-code (Terraform), mature observability (Datadog/Dynatrace/LogicMonitor), implement GitOps workflows (GitHub/GitHub Actions), script automation (Python/PowerShell), support incident response, and mentor peers to ensure scalable, secure, and auditable IT operations.
The summary above was generated by AI

Senior Infrastructure Engineer (IT Operations & Network)

Overview

Some of the world’s most innovative global enterprise software companies struggle to find technical delivery partners capable of matching their rigorous standards. These teams need a partner that can co-own complex problems from within their own IT environment.

Enter EverOps - the premier Embedded Service Provider. We partner directly with customer IT teams to assess and address mission-critical delivery and infrastructure challenges.

You’ll operate at the intersection of network infrastructure, cloud platforms, and observability, building resilient, well-monitored, and automated IT operations environments.

The Challenge

We’re hiring a Senior Infrastructure Engineer with deep mastery in networking and IT operations to lead a critical evolution of our infrastructure and monitoring environment. This role will modernize how we manage network infrastructure, mature observability practices, and drive infrastructure as code across cloud and on-prem environments.

The Mission

As a Senior Infrastructure Engineer, you will join our U.S.-Based Virtual Operating Center, working within a dynamic team to own and evolve enterprise network and IT operations infrastructure across cloud and hybrid environments. Your primary mission will focus on strengthening network architecture, maturing monitoring and observability, and building automated, infrastructure-as-code operations to improve reliability, visibility, and scalability.

You will be expected to lead by example - architecting and troubleshooting complex network environments, building out monitoring and observability stacks with tools like Datadog, Dynatrace, and LogicMonitor, and driving infrastructure as code using Terraform and GitHub-based operations, while mentoring peers and establishing best practices to ensure scalable, secure, and repeatable IT operations.

What You’ll Do
  • Lead design, implementation, and troubleshooting of enterprise network infrastructure (routing, switching, firewalls, VPN, wireless)

  • Own and mature observability strategy across infrastructure and applications using Datadog, Dynatrace, LogicMonitor, SNMP-based polling, and similar platforms

  • Build proactive monitoring, alerting, and dashboarding to drive faster detection and resolution of issues

  • Design and manage network and cloud infrastructure using Terraform (or equivalent IaC tools)

  • Develop reusable modules for network configurations, monitoring policies, and cloud infrastructure components

  • Implement version-controlled infrastructure configurations with full auditability

  • Leverage GitHub (GitOps) for:

  • Source control of infrastructure and network configurations

  • Pull request-based change management

  • CI/CD pipelines (GitHub Actions) for infrastructure deployments

  • Enforce approval workflows, testing, and promotion across environments (dev → prod)

  • Treat infrastructure changes as code with full traceability and rollback capability

  • Manage and optimize infrastructure across cloud platforms (AWS, Azure, GCP) from a networking and operations lens

  • Automate routine IT operations tasks through scripting to reduce manual effort

  • Support incident response and root cause analysis for network and infrastructure issues

  • Maintain documentation and diagrams reflecting current network and infrastructure state

You Have
  • 5+ years in IT infrastructure, network engineering, or IT operations

  • Strong networking fundamentals: routing, switching, firewalls, VPN, DNS, DHCP, wireless

  • Experience working with cloud platforms (AWS, Azure, GCP) in a networking or infrastructure capacity

  • Hands-on experience with monitoring and observability tools such as Datadog, Dynatrace, and LogicMonitor

  • Expert-level knowledge of SNMPv2 and SNMPv3, including MIBs, traps, polling architecture, and secure configuration

  • Strong scripting/automation skills (Python, PowerShell, or similar)

  • Experience with Terraform (or equivalent IaC tools) in production environments

  • Experience using GitHub (or similar) for CI/CD and infrastructure automation

  • Experience with APIs and system integrations

  • Solid understanding of IT operations best practices, incident management, and troubleshooting methodology

  • Familiarity managing hybrid environments (cloud and on-prem)

Extra Awesome
  • Experience with SolarWinds or other network performance monitoring tools

  • Experience with open-source NMS platforms (e.g., LibreNMS, Zabbix, Nagios, Cacti, Icinga)

  • Relevant networking certifications (CCNA/CCNP or equivalent)

  • Experience with cloud networking constructs (VPCs, transit gateways, load balancers, peering)

  • Experience with Slack / ITSM tools (e.g., Jira, ServiceNow, Freshservice)

  • Familiarity with security frameworks (NIST, SOC2)

  • Experience building runbooks, playbooks, or automated remediation workflows

  • Incident Response / Forensics awareness to assist with security-related investigations

Benefits
  • 100% Remote Workplace: We’ve been remote since Day 1!

  • Unlimited Paid Time Off.

  • Equity: Become a true owner of the company.

  • 401k with company contribution and sponsored healthcare.

  • Professional Growth: Access to training and certification programs to accelerate your career.

Similar Jobs

2 Days Ago
Remote
United States
165K-165K Annually
Senior level
165K-165K Annually
Senior level
Security • Cybersecurity
Design, build, and maintain scalable AWS and Azure infrastructure using Infrastructure-as-Code. Develop automation, deployment pipelines, and production release processes. Contribute code to cloud services, manage networking and IAM, ensure availability and security, and respond to incidents while collaborating on observability and security practices.
Top Skills: AnsibleAWSAzureDockerGoIamKubernetesPackerPythonTerraform
Yesterday
Remote or Hybrid
USA
140K-215K Annually
Senior level
140K-215K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Design, build, and operate large-scale LLM infrastructure and data platforms for training, fine-tuning, and inference. Provision GPU clusters, optimize GPU utilization, implement model lifecycle management, deploy inference frameworks, create evaluation and observability systems, and collaborate with data scientists to productionize AI capabilities. Mentor engineers and enforce MLOps/DataOps best practices.
Top Skills: AirflowAnsibleAWSCudaDaskDockerFlinkGCPGpuJaxKubernetesLangchainLlamaindexMegatronMlflowNvidia DriversOciPythonPyTorchRaySagemakerSlurmSparkTerraformTpuTriton Inference ServerVertex AiVllm
Yesterday
Easy Apply
Remote
USA
Easy Apply
186K-219K Annually
Senior level
186K-219K Annually
Senior level
Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Build and maintain multi-cloud infrastructure standards, self-service tooling, and paved-road defaults across AWS and GCP. Partner with compute, security, and networking teams to embed guardrails, drive standardization, lead cross-team initiatives, and resolve systemic operational issues to ensure secure, scalable, and reliable platform infrastructure.
Top Skills: AWSDnsGCPGenerative AiGoLoad BalancersPulumiPythonTerraformTransit GatewaysVpc

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account