
Job Overview
Location
USA | Remote
Job Type
Full-time
Category
Software Engineering
Date Posted
July 4, 2026
Full Job Description
📋 Description
- • Architect, build, and maintain a hybrid infrastructure platform spanning AWS cloud and on-premise bare metal data centers to support AI/ML research and production workloads.
- • Design and manage infrastructure using Infrastructure-as-Code (IaC) with Terraform to ensure environments are reproducible, version-controlled, and fully automated.
- • Deploy and operate Kubernetes clusters across both cloud and on-premise environments to provide a stable, scalable, and self-service platform for AI/ML teams.
- • Integrate and optimize Slurm job scheduler with Kubernetes to efficiently manage high-demand GPU workloads for training and inference of voice AI models.
- • Provision, configure, and lifecycle-manage bare metal servers using tools like PXE boot and MAAS for high-performance computing requirements.
- • Implement and maintain networking solutions including CNI plugins and service mesh to support low-latency, high-throughput communication across hybrid environments.
- • Design and manage storage solutions using CSI drivers and S3-compatible systems to handle massive audio data ingestion and model artifact storage.
- • Build and operate a comprehensive observability stack including monitoring, logging, and tracing to ensure platform reliability, performance, and rapid incident response.
- • Automate operational tasks, incident remediation, and performance tuning to reduce toil and improve platform resilience.
- • Collaborate directly with AI researchers and ML engineers to understand their infrastructure needs and co-build tools that accelerate model development and deployment cycles.
- • Automate the provisioning and lifecycle management of single-tenant, managed deployments for enterprise customers using self-service workflows.
- • Maintain a strong AI-first mindset by actively using and integrating advanced AI tools into daily workflows to improve efficiency and innovation.
- • Adapt rapidly to evolving AI technologies and infrastructure demands, embracing change and continuous learning as core to the role.
- • Treat infrastructure as a product, continuously improving the developer experience through automation, documentation, and user-centric design.
- • Ensure cost efficiency and performance optimization across hybrid cloud and on-premise environments using FinOps principles and resource scheduling strategies.
🎯 Requirements
- • 5+ years of experience in Platform Engineering, DevOps, or Site Reliability Engineering (SRE)
- • Proven, hands-on experience building and managing production infrastructure with Terraform
- • Expert-level knowledge of Kubernetes architecture and operations in a large-scale environment
- • Experience with high-performance compute (HPC) job schedulers, specifically Slurm, for managing GPU-intensive AI workloads
- • Experience managing bare metal infrastructure, including server provisioning (e.g., PXE boot, MAAS), configuration, and lifecycle management
- • Strong scripting and automation skills (e.g., Python, Go, Bash)
🏖️ Benefits
- • Medical, dental, vision benefits
- • Annual wellness stipend
- • Mental health support
- • Life, STD, LTD Income Insurance Plans
- • Unlimited PTO
- • Parental leave
- • Flexible schedule
- • 12 Paid US company holidays
- • Quarterly personal productivity stipend
- • One-time stipend for home office upgrades
- • 401(k) plan with company match
- • Tax Savings Programs
- • Learning / Education stipend
- • Participation in talks and conferences
- • Employee Resource Groups
- • AI enablement workshops / sessions
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About Deepgram Inc.
Deepgram builds end-to-end speech AI infrastructure that converts live or recorded audio into text and insights. The company trains large-scale neural networks on GPU clusters to deliver low-latency transcription, keyword detection, and speaker diarization through a single API. Developers use the platform for call centers, meetings, podcasts, and voice bots, paying per minute or hosting the engine on-premise. Founded in 2015 and headquartered in San Francisco, Deepgram serves enterprises seeking accurate, private, and customizable speech recognition without vendor lock-in.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

Atomic Financial Inc.
2 months ago

PermitFlow Inc.
2 months ago

ElevenLabs Inc.
2 months ago

Imprint Technologies Inc.
1 month ago
