
Job Overview
Location
Remote - USA
Job Type
Full-time
Category
Software Engineering
Date Posted
July 1, 2026
Full Job Description
đź“‹ Description
- • As a Sr. Principal Software Engineer at Cerence Inc., you will drive the future of mobility by optimizing and deploying high-performance LLM inference pipelines for AI-powered voice companions in vehicles, directly impacting the safety, connectivity, and enjoyment of drivers and passengers worldwide.
- • You will own inference runtimes across data center, edge, and embedded platforms, pushing model performance through quantization, kernel fusion, cache optimization, and latency/throughput improvements that directly impact production products, while enabling efficient, reliable deployment without external vendor dependency.
- • You will join a global, customer-centric, collaborative, and fast-paced team headquartered in Burlington, Massachusetts, with 16 offices across Europe, Asia, and North America, dedicated to advancing transportation user experiences through AI innovation and continuous learning opportunities.
- • You will deepen your expertise in vLLM, TensorRT-LLM, llama.cpp, and QAIRT, master CUDA kernel development and profiling, implement advanced quantization techniques (INT8, INT4, FP4, FP8, AWQ, GPTQ), optimize KV cache performance, and design latency-tuning strategies like batching and speculative decoding—positioning yourself as a leader in production-grade ML inference optimization for embedded automotive systems.
🎯 Requirements
- • Proven experience optimizing ML inference performance in production environments
- • Deep understanding of GPU architecture and memory hierarchies
- • Hands-on experience with CUDA and low-level performance tuning
- • Experience deploying models beyond research environments into production
- • Expertise with inference engines: vLLM, TensorRT-LLM, llama.cpp, QAIRT
- • Proficiency in CUDA kernel development and profiling
- • Knowledge of quantization techniques: INT8, INT4, FP4, FP8, AWQ, GPTQ
- • Experience with KV cache optimization and memory layout design
- • Familiarity with latency optimization strategies: batching, speculative decoding, continuous batching
🏖️ Benefits
- • Competitive salary range of $185,000.00 USD - $280,000.00 USD
- • Annual bonus opportunity
- • Comprehensive insurance coverage (medical, dental, vision, life, and disability)
- • Paid time off and paid holidays
- • Company contribution to RRSP (Registered Retirement Savings Plan)
- • Equity awards for certain positions and levels
- • Remote and/or hybrid work availability
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About Cerence Inc.
Cerence Inc. develops AI-powered mobility assistant technologies, including voice, natural language understanding, and human-machine interaction software. The company supplies automakers and mobility OEMs with embedded and cloud solutions that enable conversational in-car experiences, voice biometrics, navigation, and content access. Cerence serves global car manufacturers, tier-one suppliers, and mobility service providers, offering scalable platforms that integrate with vehicle infotainment systems and digital cockpits. Its portfolio spans speech recognition, edge AI, and cloud services designed to enhance driver safety and user experience across passenger vehicles, two-wheelers, and commercial fleets.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

Trustly Group AB
2 months ago

Trustly Group AB
2 months ago

Ruby Labs Ltd.
2 months ago

Rerun Technologies Inc.
2 months ago