
Job Overview
Location
USA | Remote
Job Type
Full-time
Category
Software Engineering
Date Posted
July 10, 2026
Full Job Description
📋 Description
- • Own end-to-end delivery of AI infrastructure programs—from model training pipelines and experiment tracking to inference serving and production monitoring
- • Define technical architecture, integration patterns, and rollout strategies for new ML systems and tooling (e.g., vector databases, model servers, evaluation frameworks, prompt engineering platforms)
- • Serve as connective tissue between ML research, ML engineering, product, and data teams to align on ML system requirements, capability roadmaps, and deployment timelines
- • Drive cost and latency optimization for real-time inference workloads at scale
- • Build lightweight internal tools and processes to accelerate ML iteration cycles (experiment tracking, model versioning, A/B testing infrastructure)
- • Identify and resolve technical bottlenecks in training pipelines, serving infrastructure, and model evaluation workflows
- • Work closely with ML practitioners to translate research breakthroughs into scalable, observable systems
🎯 Requirements
- • 5+ years of program management or technical leadership in ML infrastructure, ML platforms, or AI tooling (or equivalent)
- • Strong technical acumen in ML systems—ideally hands-on experience as an ML engineer, systems engineer, or ML infrastructure engineer
- • Experience coordinating cross-functional ML programs (e.g., model training → evaluation → serving → monitoring)
- • Proven ability to translate ML/research requirements into robust, scalable infrastructure
🏖️ Benefits
- • Competitive salary
- • Comprehensive benefits package
- • Flexible work arrangements
- • Opportunities for professional growth and development
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About Deepgram Inc.
Deepgram builds end-to-end speech AI infrastructure that converts live or recorded audio into text and insights. The company trains large-scale neural networks on GPU clusters to deliver low-latency transcription, keyword detection, and speaker diarization through a single API. Developers use the platform for call centers, meetings, podcasts, and voice bots, paying per minute or hosting the engine on-premise. Founded in 2015 and headquartered in San Francisco, Deepgram serves enterprises seeking accurate, private, and customizable speech recognition without vendor lock-in.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

Atomic Financial Inc.
2 months ago

PermitFlow Inc.
2 months ago

ElevenLabs Inc.
2 months ago

Imprint Technologies Inc.
1 month ago
