STN Inc

United States
99 Total Employees
Year Founded: 2016

Jobs at STN Inc

Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.

Recently posted jobs

11 Hours AgoSaved
Remote
USA
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
Designs, deploys, and operates high-performance InfiniBand and RoCE networking fabrics for GPU clusters, customer connectivity, and multi-site WANs. Manages routing, switching, BGP peering, public IPs, DDoS protection, automation, monitoring, troubleshooting, documentation, capacity planning, and security audits. Partners with the NOC and supports customer network architectures across layers 1 through 7.
11 Hours AgoSaved
Remote
USA
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
Owns the hardware lifecycle for GPU and infrastructure assets, including fleet health monitoring, vendor RMA workflows, firmware and BIOS upgrades, burn-in testing, failure investigation, inventory accuracy, capacity planning, spare-parts strategy, runbook creation, and new platform qualification. The role serves as the technical owner of physical compute platforms and supports high-density GPU infrastructure across sites.
11 Hours AgoSaved
Remote
USA
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
Build and operate the multi-tenant platform layer for a GPU cloud service. Responsibilities include Kubernetes-based orchestration, tenant isolation, platform APIs and tools, GPU cluster automation, infrastructure as code, reliability and observability partnerships, customer issue support, environment templates, rollout tooling, architecture reviews, and reducing operational toil.
11 Hours AgoSaved
Remote
USA
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
Owns reliability, observability, and incident response for a GPUaaS platform. Defines SLOs, builds monitoring and alerting systems, leads major incidents and post-incident reviews, automates operational processes, maintains runbooks, manages on-call operations, coordinates with engineering teams, drives chaos testing, reports SLA performance, and mentors junior engineers.
11 Hours AgoSaved
Remote
USA
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
Architect, deploy, optimize, and operate large-scale GPU clusters for AI training and inference. Responsibilities include distributed PyTorch and NCCL tuning, GPU networking and storage optimization, scheduling, benchmarking, troubleshooting, automation, monitoring, and performance engineering across compute, networking, storage, and software layers. The role supports production environments with hundreds to thousands of GPUs and partners with ML engineers to improve training scalability and inference efficiency.
11 Hours AgoSaved
Remote
USA
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
Owns security operations and compliance for a multi-tenant GPUaaS platform. Maintains SOC 2 and SOC 3 programs, coordinates audits and penetration tests, manages IAM and security tooling, leads vulnerability management and incident response, supports customer security reviews and contracts, maintains policies, and drives security awareness and phishing programs.
11 Hours AgoSaved
Remote
USA
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
Monitor GPUaaS infrastructure, triage alerts, execute runbooks, coordinate incident response, communicate with customers, manage maintenance windows, resolve Tier 1 tickets, and improve monitoring and operational procedures. The role provides 24/7 rotating coverage, including nights, weekends, and holidays, while maintaining shift handoffs and active-incident documentation.
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
Provides Tier 1 and Tier 2 technical support for GPUaaS customers through ticketing, email, and chat. Responsibilities include triaging issues against SLAs, resolving common problems using Linux and networking knowledge, escalating complex incidents, maintaining ticket records, creating documentation, supporting account lifecycle changes, participating in rotating on-call coverage, and improving runbooks and customer satisfaction.