
Job Overview
Location
USA | Remote
Job Type
Full-time
Category
Software Engineering
Date Posted
July 9, 2026
Full Job Description
📋 Description
- • Take Deepgram's Speech and Conversational models and get them running on embedded and low-power consumer hardware — defining the architecture for on-device, real-time inference across a diverse range of processors and accelerators.
- • Optimize models for constrained targets through quantization, pruning, distillation, operator fusion, and architecture-specific compilation to meet strict latency, memory, power, and thermal budgets.
- • Write and optimize performance-critical runtime code (C, C++, and/or Rust) for embedded environments, including bare-metal and real-time operating systems such as FreeRTOS and Zephyr.
- • Integrate with industry-standard edge inference runtimes and vendor NPU/DSP toolchains, mapping model graphs efficiently onto on-device accelerators and CPU/GPU/NPU heterogeneity.
- • Build the on-device runtime plumbing: model packaging, deployment pipelines, over-the-air update mechanisms, and lightweight telemetry for devices operating with limited or intermittent connectivity.
- • Establish repeatable benchmarking and validation across target hardware — measuring latency, accuracy, power consumption, memory footprint, and resource utilization — and catch regressions before they ship.
- • Partner with silicon and device vendors on SDK integration and performance tuning, getting our models to run efficiently on new chipsets and reference platforms.
- • Collaborate with Research and Engine teams to influence model architectures toward edge-friendly designs from the start, reducing the optimization burden at deployment time.
🎯 Requirements
- • Experience delivering production systems on resource-constrained hardware — embedded systems, mobile, edge AI, or small low-power devices.
- • Strong proficiency in C, C++, and/or Rust, with experience writing performance-critical code for constrained environments.
- • Hands-on experience with model optimization for on-device deployment, including quantization, pruning, knowledge distillation, or architecture-specific compilation.
🏖️ Benefits
- • Competitive salary
- • Comprehensive benefits package
- • Flexible work arrangements
- • Professional development opportunities
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About Deepgram Inc.
Deepgram builds end-to-end speech AI infrastructure that converts live or recorded audio into text and insights. The company trains large-scale neural networks on GPU clusters to deliver low-latency transcription, keyword detection, and speaker diarization through a single API. Developers use the platform for call centers, meetings, podcasts, and voice bots, paying per minute or hosting the engine on-premise. Founded in 2015 and headquartered in San Francisco, Deepgram serves enterprises seeking accurate, private, and customizable speech recognition without vendor lock-in.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

Atomic Financial Inc.
2 months ago

PermitFlow Inc.
2 months ago

ElevenLabs Inc.
2 months ago

Imprint Technologies Inc.
1 month ago
