Deepgram Inc. logo

Embedded AI Engineer, On-Device Models

Job Overview

Location

USA | Remote

Job Type

Full-time

Category

Software Engineering

Date Posted

July 9, 2026

Full Job Description

📋 Description

  • Take Deepgram's Speech and Conversational models and get them running on embedded and low-power consumer hardware — defining the architecture for on-device, real-time inference across a diverse range of processors and accelerators.
  • Optimize models for constrained targets through quantization, pruning, distillation, operator fusion, and architecture-specific compilation to meet strict latency, memory, power, and thermal budgets.
  • Write and optimize performance-critical runtime code (C, C++, and/or Rust) for embedded environments, including bare-metal and real-time operating systems such as FreeRTOS and Zephyr.
  • Integrate with industry-standard edge inference runtimes and vendor NPU/DSP toolchains, mapping model graphs efficiently onto on-device accelerators and CPU/GPU/NPU heterogeneity.
  • Build the on-device runtime plumbing: model packaging, deployment pipelines, over-the-air update mechanisms, and lightweight telemetry for devices operating with limited or intermittent connectivity.
  • Establish repeatable benchmarking and validation across target hardware — measuring latency, accuracy, power consumption, memory footprint, and resource utilization — and catch regressions before they ship.
  • Partner with silicon and device vendors on SDK integration and performance tuning, getting our models to run efficiently on new chipsets and reference platforms.
  • Collaborate with Research and Engine teams to influence model architectures toward edge-friendly designs from the start, reducing the optimization burden at deployment time.

🎯 Requirements

  • Experience delivering production systems on resource-constrained hardware — embedded systems, mobile, edge AI, or small low-power devices.
  • Strong proficiency in C, C++, and/or Rust, with experience writing performance-critical code for constrained environments.
  • Hands-on experience with model optimization for on-device deployment, including quantization, pruning, knowledge distillation, or architecture-specific compilation.

🏖️ Benefits

  • Competitive salary
  • Comprehensive benefits package
  • Flexible work arrangements
  • Professional development opportunities

Skills & Technologies

Go
Rust
Linux
Data Science
Remote

Ready to Apply?

You will be redirected to an external site to apply.

AI Job Fit Analysis
Pro

See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.

Deepgram Inc. logo
Deepgram Inc.
Visit Website

About Deepgram Inc.

Deepgram builds end-to-end speech AI infrastructure that converts live or recorded audio into text and insights. The company trains large-scale neural networks on GPU clusters to deliver low-latency transcription, keyword detection, and speaker diarization through a single API. Developers use the platform for call centers, meetings, podcasts, and voice bots, paying per minute or hosting the engine on-premise. Founded in 2015 and headquartered in San Francisco, Deepgram serves enterprises seeking accurate, private, and customizable speech recognition without vendor lock-in.

Get more remote jobs like this

Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.

Newsletter

Weekly remote jobs and featured talent.

No spam. Only curated remote roles and product updates. You can unsubscribe anytime.

Similar Opportunities

Expired
Atomic Financial Inc. logo

Atomic Financial Inc.

Remote
Full-time
Expired Jul 15, 2026
Grafana
OAuth
Remote
+1 more

2 months ago

Expired
PermitFlow Inc. logo

PermitFlow Inc.

New York City, NY
Full-time
Expired Jul 15, 2026
Hybrid

2 months ago

Expired
UAE
Full-time
Expired Jul 15, 2026
Python
Mid-level
Remote

2 months ago

Imprint Technologies Inc. logo

Imprint Technologies Inc.

Remote
Full-time
Expires Aug 12, 2026
Rails
Remote

1 month ago