R

Embedded AI Engineer

Remotive
2 hours ago
Full-time
On-site
East New York, New York, United States
Role Description

Deepgram's speech AI models are among the fastest and most accurate in the world — and the next wave of voice experiences won't live only in the cloud. They'll run directly on the small, low-power devices people carry, wear, and keep around their homes: phones, earbuds, wearables, appliances, cameras, and purpose-built consumer hardware. Putting state-of-the-art speech models on devices with tight memory, compute, thermal, and battery budgets is a fundamentally different engineering problem, and it's one of the most important frontiers for bringing voice AI to everyone.

As an

Embedded AI Engineer

, you will take Deepgram's models and make them run — fast, accurately, and efficiently — on resource-constrained embedded and edge platforms. You'll work across the stack:

Optimizing and compiling models for on-device inference.

Writing performance-critical runtime code.

Squeezing every last millisecond and milliwatt out of a wide range of mobile application processors, embedded SoCs, microcontrollers, and dedicated AI accelerators.

Your work directly enables a new class of private, offline-capable, real-time voice experiences on the devices closest to the user.

This role is a great fit whether you're a hands-on senior embedded engineer who wants to go deep on a hard problem, or a staff-level technical leader who wants to define how Deepgram's voice AI gets onto consumer hardware and raise the bar for the engineers around you. We'll set the level to your experience.

Qualifications

Experience delivering production systems on resource-constrained hardware — embedded systems, mobile, edge AI, or small low-power devices.

Strong proficiency in C, C++, and/or Rust, with experience writing performance-critical code for constrained environments.

Hands-on experience with model optimization for on-device deployment, including quantization, pruning, knowledge distillation, or architecture-specific compilation.

Familiarity with edge inference runtimes (e.g., ONNX Runtime, TensorRT, TFLite, ExecuTorch) and/or vendor-specific NPU/DSP toolchains.

A strong understanding of hardware-software interaction — CPU/GPU/NPU/DSP architectures, memory hierarchies, fixed-point/integer arithmetic, and power management — and how they affect inference performance.

Experience working close to the metal: bare-metal or RTOS environments (e.g., FreeRTOS, Zephyr), embedded Linux, or microcontroller and edge SoC development.

Strong communication skills and a builder mindset — you can scope an ambiguous optimization problem, drive it to a measurable result, and explain the tradeoffs clearly.

Requirements

Find deep satisfaction in making a large model run on a tiny device — and still hit accuracy and latency targets.

Want to work at the intersection of AI and hardware, where optimization isn't optional but existential.

Are energized by the back-and-forth of getting a model to sing on a new chipset, runtime, or accelerator.

Believe on-device AI is the next major deployment frontier and want to define how speech AI gets there for consumers.

Prefer hard, constrained, ship-it problems over open-ended research — you want to see your work running in people's hands.

Care about the details that don't show up in a cloud benchmark: cold-start time, power draw, thermals, and memory fragmentation.

Benefits

Competitive salary and equity options.

Comprehensive health benefits.

Flexible work hours and remote work options.

Opportunities for professional development and growth.