This job has expired
This position was posted on April 24, 2026 and is likely no longer accepting applications. We've kept it here for historical reference. Check out the similar jobs below!

Job Overview
Location
San Francisco
Job Type
Full-time
Category
Software Engineering
Date Posted
April 24, 2026
Full Job Description
đź“‹ Description
- • As a Software Engineer - Voice AI (Inference Runtime) at BaseTen Inc., you will be the primary owner of the company's in-house inference stack for Voice AI models, playing a critical role in enabling cutting-edge voice technologies for customers across industries such as productivity, customer service, healthcare, and education.
- • Your day-to-day responsibilities include owning and leading Voice AI product areas end-to-end—from architecture and system design through implementation, rollout, and long-term production operations—designing, building, and operating real-time, large-scale, high-performance model serving systems for STT, TTS, and voice agent workloads with clear SLOs, driving cross-team collaboration with sister engineering teams to solve full-stack technical problems, and mentoring teammates through code reviews, design docs, and technical leadership.
- • You will join a small founding team focused on bringing state-of-the-art open-source voice models into production, collaborating closely with Forward Deployed Engineers, Model Performance Engineers, and the Core Product and Training Platform teams to push the boundaries of Voice AI and enable self-serve adoption through ergonomic APIs and SDKs.
- • This role offers the opportunity to make a meaningful impact on people’s daily lives by helping reshape industries through voice technology, while gaining deep expertise in ML infrastructure, real-time systems, and developer platform design at a fast-growing AI company backed by top-tier investors.
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About BaseTen Inc.
BaseTen provides a serverless, GPU-accelerated platform that lets machine-learning teams deploy, scale and monitor custom models behind autoscaling inference endpoints. The service abstracts infrastructure management, supports PyTorch, TensorFlow and Hugging Face artifacts, and offers built-in observability, A/B testing and fine-tuning. Customers integrate via REST or GraphQL APIs and pay only for compute used. Founded in 2019 and headquartered in San Francisco, BaseTen targets data scientists and product teams seeking production-grade ML serving without Kubernetes complexity.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

xAI Corporation
3 months ago

Ramp Business Corporation
2 months ago

Hims & Hers Health, Inc.
2 months ago

Known Holdings, Inc.
3 months ago