Top Tech Jobs & Startup Jobs in Boston, MA

YesterdaySaved
Remote
USA
Senior level
Senior level
Software
Build and own a multi-tenant web platform for managing AI infrastructure. Develop Next.js and React features, server-side authentication and API proxying, generated TypeScript SDKs, resilient data synchronization, design-system components, accessibility, performance, testing, observability, and production deployment. Collaborate with backend and globally distributed teams, review code, troubleshoot production issues, and mentor junior engineers.
Top Skills: Base UiCvaDesign TokensDockerEslintFigmaGithub ActionsGoJoseJwtKubernetesLadleMswNext.JsNode.Js 22+Oauth2OidcOpenapiOpentelemetryPlaywrightPnpmPosthogPrettierRadixReact 19SentryStorybookTailwind Css V4Tanstack QueryTesting LibraryTurborepoTypescriptVitestZod
YesterdaySaved
Remote
USA
Senior level
Senior level
Software
Set technical direction for k0rdent across core Kubernetes management, AI infrastructure, AI applications, and observability. Design cross-cutting distributed systems, define architectural standards, write code, resolve complex technical issues, mentor senior engineers, guide multi-quarter strategy, collaborate with product and engineering leadership, and represent the company with customers and open-source communities.
Top Skills: Ai/Ml RuntimesAutoscalingCncfDistributed SystemsGpu SchedulingK0RdentKcmKsmKubernetesLogsMetricsModel ServingOpentelemetrySlosTraces
Senior level
Software
Design and implement infrastructure services for a GPU-as-a-Service platform. Build REST and gRPC APIs, bare-metal provisioning workflows, Kubernetes cluster automation, reconciliation loops, and durable asynchronous operations. Manage server enrollment, inspection, OS provisioning, cluster creation, hardware inventory interfaces, error handling, idempotency, and multi-tenant infrastructure reliability.
Top Skills: ArgocdBmcCluster ApiDpusFluxGitopsGoGpusGrpcIpxeK0RdentKubernetesMetal3NicsOpentofuPxeRedfishRestTemporalTerraform
Senior level
Software
Design and develop Rust and Go control-plane microservices, Kubernetes operators, reconciliation engines, and APIs for bare-metal AI infrastructure. Integrate datacenter networking systems, DPUs, network operating systems, and hardware-management platforms. Build observability with OpenTelemetry, troubleshoot distributed systems, and maintain rigorous testing practices. The role requires expertise in datacenter networking protocols, Linux networking, Kubernetes, provisioning, security, virtualization, and high-performance interconnects, along with architectural documentation and open-source collaboration.
Top Skills: AxumBgpBiosBluefield DpuCalicoCiliumController-RuntimeCumulus LinuxDpdkDpfEvpnGoGrpcHost-Based NetworkingInfinibandInitramfsIpmiIpxeJwtKeycloakKmsKube-RsKubernetesKubernetes OperatorsKubevirtKvmL3VniLinuxLinux NetlinkMp-BgpNmx-MNvidia DocaNvlinkNvueOauth2OpentelemetryOvsPkiPostgresProtobufPxeRbacRedfishRestRocev2RustSecure BootSonicSpiffe/SvidSQLSqlxSystemdTlsTokioTonicTpmUefiVaultVxlanX.509
8 Days AgoSaved
Remote
USA
Senior level
Senior level
Software
Design and build enterprise LLM-serving infrastructure on Kubernetes, including GPU scheduling, scaling, model lifecycle management, Helm packaging, offline deployments, platform integrations, and production observability. Contribute across a multi-service Go codebase, create design documentation, participate in reviews, and help guide engineering direction within an autonomous remote-first team.
Top Skills: Api GatewaysCi/CdClaude CodeGoGpu OrchestrationGpu TelemetryHelmInfrastructure As CodeKubernetesLlm InferenceOidcOpenai CodexReactSsoTypescript
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Senior level
Software
Own the vision, roadmap, backlog, and priorities for k0rdent AI observability across GPU infrastructure and distributed AI workloads. Define integrations for metrics, tracing, logs, telemetry, and alerting across Kubernetes, networking, storage, schedulers, inference platforms, and databases. Partner with engineering, marketing, field teams, customers, and ecosystem vendors to shape requirements, evaluate trade-offs, develop positioning, and support deployments.
Top Skills: DcgmDdnDpuElasticsearchInfinibandJaegerKubernetesLokiMilvusNvidia BluefieldOpensearchOpentelemetryPrometheusQdrantRocev2SlurmSmartnicSr-IovTempoTensorrt-LlmTritonVast DataVllmWeka
12 Days AgoSaved
Remote
USA
Senior level
Senior level
Software
Deploys, integrates, operates, and tunes high-performance NFS storage for Kubernetes, GPU, and AI workloads. Responsibilities include CSI integration, Linux and network optimization, storage lifecycle management, air-gapped platform support, infrastructure-as-code, GitOps automation, observability, and troubleshooting distributed storage performance and reliability across hybrid, edge, and bare-metal environments.
Top Skills: Argo CdCluster ApiContainer Storage Interface (Csi)DdnDell PowerscaleFluxGitopsHarborK0RdentKubernetesLinuxNfsOpentofuPki/TlsRdmaTerraformVastWeka
12 Days AgoSaved
Remote
USA
Senior level
Senior level
Software
Deploy, integrate, operate, and optimize high-performance NFS storage for GPU-accelerated Kubernetes and AI platforms. Configure Linux storage and networking, integrate CSI storage into k0s and K0rdent clusters, automate provisioning through Terraform/OpenTofu and GitOps, and build observability for capacity, performance, and reliability. Diagnose end-to-end storage issues across hybrid, edge, and air-gapped environments while establishing operational standards.
Top Skills: ArgocdCapiCluster ApiCsiDell PowerscaleFluxGitopsHarborK0RdentK0SKubernetesLinuxNfsOpentofuPkiRdmaTerraformTlsVast
27 Days AgoSaved
Remote
USA
Senior level
Senior level
Software
Design and build Go-based control-plane services, CSI drivers, storage integrations, observability tools, and infrastructure automation for Kubernetes AI platforms. Integrate enterprise and scale-out storage, support bare-metal and hybrid provisioning, and deliver declarative Terraform/OpenTofu and GitOps workflows. Architect reliable storage solutions for Kubernetes clusters, air-gapped environments, and sovereign infrastructure, including artifact mirrors and PKI/TLS.
Top Skills: ArgocdCluster ApiDell PowerscaleFluxGoGpudirect StorageHarborK0RdentK0SKubernetesKubernetes CsiLinuxNfsOpentofuPkiPythonTerraformTlsVast
27 Days AgoSaved
Remote
USA
Senior level
Senior level
Software
Own the vision, roadmap, backlog, and priorities for k0rdent AI storage across GPU clusters. Define capabilities for parallel file systems, object and block storage, tiering, snapshots, data protection, and GPU-optimized data paths. Translate customer requirements into product direction, partner with engineering on scalable storage features, develop positioning and field assets, support strategic accounts, and represent Mirantis with customers, analysts, and ecosystem partners.
Top Skills: BeegfsContainer Storage Interface (Csi)DaosGpfs/Spectrum ScaleGpudirect StorageK0Rdent AiKubernetesLinuxLustreNvmeNvme-OfRdmaS3-Compatible Object StorageSoftware-Defined StorageVastWeka
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account