
Job Overview
Location
San Francisco
Job Type
Full-time
Category
Software Engineering
Date Posted
July 10, 2026
Full Job Description
đź“‹ Description
- • Develop a first-class observability and root-cause analysis system for GPU fabrics.
- • Collect high-volume signals from switches, hosts, active probes, and inference services.
- • Reduce and correlate them in real time.
- • Understand topology and service ownership.
- • Produce actionable diagnosis while an incident is still unfolding.
- • Build a real-time telemetry engine for high-cardinality fabric, host, GPU, and workload telemetry.
- • Create service-aware fabric diagnosis collectors, probes, and topology-aware correlation.
- • Tie network behavior to RDMA operations, GPUDirect paths, KV cache transfers, prefill/decode disaggregation, and request latency.
- • Model topology, flow paths, service ownership, and failure domains.
- • Separate true fabric faults from host, NIC, GPU, kernel, driver, RDMA, scheduler, and workload failures.
- • Create clear operator workflows for triage, remediation, and post-incident learning.
🎯 Requirements
- • Staff-level or senior staff-level experience building production infrastructure software.
- • Strong distributed systems background, especially streaming systems, telemetry pipelines, diagnostics, or control-plane software.
- • Experience building systems that process high-volume, high-cardinality, noisy operational data.
- • Understanding of networking fundamentals and high-performance networks.
- • Ability to work with low-level infrastructure signals and build practical correlation, anomaly detection, or root-cause analysis systems.
🏖️ Benefits
- • Competitive compensation, including meaningful equity.
- • 100% coverage of medical, dental, and vision insurance for employee and dependents.
- • Flexible PTO policy including company-wide Winter Break.
- • Paid parental leave.
- • Fertility and family-building stipend through Carrot.
- • Company-facilitated 401(k).
- • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About BaseTen Inc.
BaseTen provides a serverless, GPU-accelerated platform that lets machine-learning teams deploy, scale and monitor custom models behind autoscaling inference endpoints. The service abstracts infrastructure management, supports PyTorch, TensorFlow and Hugging Face artifacts, and offers built-in observability, A/B testing and fine-tuning. Customers integrate via REST or GraphQL APIs and pay only for compute used. Founded in 2019 and headquartered in San Francisco, BaseTen targets data scientists and product teams seeking production-grade ML serving without Kubernetes complexity.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

Harvey AI Inc.
1 month ago

Heidi Health Pty Ltd
3 months ago
1 month ago
1 month ago
