TensorWave Logo

TensorWave

Principal Network Engineer

Posted 2 Months Ago
Remote
Hiring Remotely in USA
Expert/Leader
Remote
Hiring Remotely in USA
Expert/Leader
Design, own, and operate front-end network architecture for large-scale AI/GPU infrastructure including DCI, edge, ingress/egress and control-plane. Architect scalable Ethernet, routing, segmentation, and high-availability designs; lead deployments and troubleshooting in new data centers; maintain reference architectures and carrier relationships; collaborate on Kubernetes-centric network solutions and interface with RDMA back-end fabrics.
The summary above was generated by AI

About TensorWave

Our mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure.

 

About the Role

We’re seeking a Principal Network Engineer (L7) focused on owning and evolving large-scale, RoCEv2 data center networks powering next generation AI and ML infrastructure.
You’ll work closely with our network architect and infrastructure leadership to define how the network is designed, implemented, and operated at scale, keeping over 8,000 GPUs burring today and scaling to cluster sizes reaching over 100,000 GPUs. You will be responsible for the architectural decisions that determine performance, reliability, and operational sanity at scale.


You’ll remain hands-on with high-speed optics, switching, routing, and congestion management in production clusters, while also setting the standards, patterns, and tooling other engineers build and operate against.

 

What You’ll Do

  • As a Principal Engineer, you own end-to-end network architecture, make high-impact design decisions, and set technical direction across teams, with clear examples of systems you’ve defined and scaled

  • Define, evolve, and standardize large-scale RoCEv2 data center networks supporting AI and ML clusters from thousands to 100,000+ GPUs

  • Set and validate congestion management strategy across RDMA fabrics, including PFC, ECN, and DCQCN, based on real production behavior

  • Establish automation, validation, and observability patterns that prevent misconfiguration and eliminate manual operational work

  • Act as the technical escalation point for complex failures, scaling limits, and architectural tradeoffs in always-on, multi-tenant environments

 

Essential Skills & Qualifications

  • Bachelor’s degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience

  • Deep experience designing and operating RDMA and RoCEv2 networks in large-scale production environments supporting AI or HPC workloads

  • Expert-level knowledge of switching hardware and their NOS, such as Arista, Juniper, and custom solutions using SONiC, including high-speed Ethernet fabrics

  • Proven hands-on experience with congestion management and performance tuning using PFC, ECN, and DCQCN

  • Strong experience with high-speed optics and cabling including 400G, 800G, and AEC, AOC, DAC, and structured cabling at scale

  • Strong automation mindset, with experience using Python, Ansible, Terraform, Git, and production observability tooling

  • We’re looking for engineers who operate comfortably at scale, make hard calls with incomplete data, and take responsibility for systems that must work under sustained load. The solutions that work on a handful of devices will not work at Exascale.

 

What We Offer

  • Stock Options

  • 100% paid Medical, Dental, and Vision insurance for Employees

  • Company Health Savings Account Contributions

  • 100% paid Short Term and Long Term Disability Insurance for Employees

  • Life and Voluntary Supplemental Insurance Options

  • Other Insurance Options, such as Pet & Legal Insurance

  • Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support

  • Flexible Spending Account

  • 401(k)

  • Employee Assistance Program

  • Flexible PTO

  • Paid Holidays

  • Parental Leave

  • Other In-Office Perks

 

Equal Employment Opportunity

TensorWave is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of any protected status under applicable law.

 

Reasonable Accommodations

TensorWave provides reasonable accommodations in accordance with applicable laws. If you require accommodation during the hiring process, please contact [email protected].

 

Employment Eligibility

All offers of employment are contingent upon verification of identity and authorization to work in the United States, as required by law.

 

Background Checks

Where permitted by law, employment may be contingent upon the successful completion of a job-related background check.

 

Data Privacy Notice

By submitting an application, you acknowledge that TensorWave may collect, use, and retain your personal information for recruiting and employment-related purposes in accordance with applicable data privacy laws.

Similar Jobs

3 Days Ago
In-Office or Remote
New Jersey, USA
175K-412K Annually
Expert/Leader
175K-412K Annually
Expert/Leader
Artificial Intelligence • Cloud • Information Technology • Consulting
Leads technical pre-sales engineering for service provider networking opportunities. Advises customers and account teams on routing, automation, security, data center, and WAN solutions; designs architectures; delivers technical presentations and demonstrations; supports proofs of concept and RFP responses; and builds executive and engineering stakeholder relationships. The role also represents the company at industry events and requires extensive expertise in service provider networking technologies, automation, and network management.
Top Skills: AnsibleApstraAristaBgpBngCarrier EthernetDockerEvpnGrpcIs-IsJerichoKubernetesL2/L3 VpnLinuxMistMplsMulticastMx SeriesNetconfNexusOpenconfigOspfParagonPtx SeriesPythonQfxRest ApisRsvpSr-TeSrv6SrxTomahawkTridentYang
4 Days Ago
Remote or Hybrid
United States
123K-185K Annually
Senior level
123K-185K Annually
Senior level
Digital Media • Fintech • Information Technology • Machine Learning • Financial Services • Cybersecurity • Automation
Architects, deploys, and manages secure automated networking across AWS, Azure, and hybrid-cloud environments. Designs Spine-Leaf networks, optimizes Direct Connect and ExpressRoute connectivity, leads infrastructure-as-code practices with Terraform and Ansible, and drives multi-cloud network product vision. The role troubleshoots network and security issues, supports tier 3 on-call operations, collaborates with stakeholders, and creates technical documentation and knowledge-sharing presentations.
Top Skills: AnsibleAviatrixAWSAws CloudwatchAws Direct ConnectAzureAzure ExpressrouteAzure MonitorCloudtrailHarnessIp NetworkingIpsecJIRANetmonRoutingSpine-Leaf ArchitectureSplunkSwitchingTerraformTracertVpc Flow LogsVrfsWireshark
8 Days Ago
Remote
Pearl Harbor, HI, USA
124K-161K Annually
Expert/Leader
124K-161K Annually
Expert/Leader
Aerospace • Information Technology • Professional Services • Security • Software
Manages and secures systems infrastructure across multiple operating systems, maintaining system integrity, documentation, policies, hardware, software, and backup recovery. Supports service deployments, vulnerability remediation, patching, STIG compliance, and user requests. Performs occasional network, information assurance, logistics, and data transport duties. Requires advanced Windows administration, scripting, virtualization, and security knowledge, with possible after-hours, weekend, holiday, and international support responsibilities.
Top Skills: AcasActive DirectoryCcnp EnterpriseDisa StigExchangeGroup PolicyIavaItil V4Jncip-EntMicrosoft Server 2016Routing And SwitchingSccmScripting LanguagesServer StigsSharepointSQLTenableVmware Horizon VdiVmware VcenterVmware VdiWindows 10Windows PowershellWsus

What you need to know about the Boston Tech Scene

Boston is a powerhouse for technology innovation thanks to world-class research universities like MIT and Harvard and a robust pipeline of venture capital investment. Host to the first telephone call and one of the first general-purpose computers ever put into use, Boston is now a hub for biotechnology, robotics and artificial intelligence — though it’s also home to several B2B software giants. So it’s no surprise that the city consistently ranks among the greatest startup ecosystems in the world.

Key Facts About Boston Tech

  • Number of Tech Workers: 269,000; 9.4% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Thermo Fisher Scientific, Toast, Klaviyo, HubSpot, DraftKings
  • Key Industries: Artificial intelligence, biotechnology, robotics, software, aerospace
  • Funding Landscape: $15.7 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Summit Partners, Volition Capital, Bain Capital Ventures, MassVentures, Highland Capital Partners
  • Research Centers and Universities: MIT, Harvard University, Boston College, Tufts University, Boston University, Northeastern University, Smithsonian Astrophysical Observatory, National Bureau of Economic Research, Broad Institute, Lowell Center for Space Science & Technology, National Emerging Infectious Diseases Laboratories

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account