Deepgram Inc. logo

Site Reliability Engineer - AI & ML Infrastructure (Kubernetes, AWS & Terraform)

Job Overview

Location

USA | Remote

Job Type

Full-time

Category

Software Engineering

Date Posted

July 4, 2026

Full Job Description

📋 Description

  • Architect, build, and maintain a hybrid infrastructure platform spanning AWS cloud and on-premise bare metal data centers to support AI/ML research and production workloads.
  • Design and manage infrastructure using Infrastructure-as-Code (IaC) with Terraform to ensure environments are reproducible, version-controlled, and fully automated.
  • Deploy and operate Kubernetes clusters across both cloud and on-premise environments to provide a stable, scalable, and self-service platform for AI/ML teams.
  • Integrate and optimize Slurm job scheduler with Kubernetes to efficiently manage high-demand GPU workloads for training and inference of voice AI models.
  • Provision, configure, and lifecycle-manage bare metal servers using tools like PXE boot and MAAS for high-performance computing requirements.
  • Implement and maintain networking solutions including CNI plugins and service mesh to support low-latency, high-throughput communication across hybrid environments.
  • Design and manage storage solutions using CSI drivers and S3-compatible systems to handle massive audio data ingestion and model artifact storage.
  • Build and operate a comprehensive observability stack including monitoring, logging, and tracing to ensure platform reliability, performance, and rapid incident response.
  • Automate operational tasks, incident remediation, and performance tuning to reduce toil and improve platform resilience.
  • Collaborate directly with AI researchers and ML engineers to understand their infrastructure needs and co-build tools that accelerate model development and deployment cycles.
  • Automate the provisioning and lifecycle management of single-tenant, managed deployments for enterprise customers using self-service workflows.
  • Maintain a strong AI-first mindset by actively using and integrating advanced AI tools into daily workflows to improve efficiency and innovation.
  • Adapt rapidly to evolving AI technologies and infrastructure demands, embracing change and continuous learning as core to the role.
  • Treat infrastructure as a product, continuously improving the developer experience through automation, documentation, and user-centric design.
  • Ensure cost efficiency and performance optimization across hybrid cloud and on-premise environments using FinOps principles and resource scheduling strategies.

🎯 Requirements

  • 5+ years of experience in Platform Engineering, DevOps, or Site Reliability Engineering (SRE)
  • Proven, hands-on experience building and managing production infrastructure with Terraform
  • Expert-level knowledge of Kubernetes architecture and operations in a large-scale environment
  • Experience with high-performance compute (HPC) job schedulers, specifically Slurm, for managing GPU-intensive AI workloads
  • Experience managing bare metal infrastructure, including server provisioning (e.g., PXE boot, MAAS), configuration, and lifecycle management
  • Strong scripting and automation skills (e.g., Python, Go, Bash)

🏖️ Benefits

  • Medical, dental, vision benefits
  • Annual wellness stipend
  • Mental health support
  • Life, STD, LTD Income Insurance Plans
  • Unlimited PTO
  • Parental leave
  • Flexible schedule
  • 12 Paid US company holidays
  • Quarterly personal productivity stipend
  • One-time stipend for home office upgrades
  • 401(k) plan with company match
  • Tax Savings Programs
  • Learning / Education stipend
  • Participation in talks and conferences
  • Employee Resource Groups
  • AI enablement workshops / sessions

Skills & Technologies

Python
AWS
Kubernetes
Terraform
Jenkins
DevOps
Remote

Ready to Apply?

You will be redirected to an external site to apply.

AI Job Fit Analysis
Pro

See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.

Deepgram Inc. logo
Deepgram Inc.
Visit Website

About Deepgram Inc.

Deepgram builds end-to-end speech AI infrastructure that converts live or recorded audio into text and insights. The company trains large-scale neural networks on GPU clusters to deliver low-latency transcription, keyword detection, and speaker diarization through a single API. Developers use the platform for call centers, meetings, podcasts, and voice bots, paying per minute or hosting the engine on-premise. Founded in 2015 and headquartered in San Francisco, Deepgram serves enterprises seeking accurate, private, and customizable speech recognition without vendor lock-in.

Get more remote jobs like this

Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.

Newsletter

Weekly remote jobs and featured talent.

No spam. Only curated remote roles and product updates. You can unsubscribe anytime.

Similar Opportunities

Expired
Atomic Financial Inc. logo

Atomic Financial Inc.

Remote
Full-time
Expired Jul 15, 2026
Grafana
OAuth
Remote
+1 more

2 months ago

Expired
PermitFlow Inc. logo

PermitFlow Inc.

New York City, NY
Full-time
Expired Jul 15, 2026
Hybrid

2 months ago

Expired
UAE
Full-time
Expired Jul 15, 2026
Python
Mid-level
Remote

2 months ago

Imprint Technologies Inc. logo

Imprint Technologies Inc.

Remote
Full-time
Expires Aug 12, 2026
Rails
Remote

1 month ago