01Overview
About the Role
We are seeking an Edge AI Research Scientist to develop next-generation speech and audio AI systems that run efficiently on smartphones, wearables, and other resource-constrained devices.
This role sits at the intersection of machine learning research, systems optimization, and production engineering. You will design compact model architectures, develop advanced compression techniques, and optimize inference pipelines that enable real-time speech AI experiences directly on end-user devices.
You will work across a broad range of voice technologies, including automatic speech recognition (ASR), text-to-speech (TTS), speech translation, speech-to-speech systems, and neural audio codecs.
The ideal candidate combines strong research credentials with hands-on implementation skills and has a deep understanding of efficient deep learning, model optimization, and hardware-aware machine learning.
What You'll Do
Research and Model Development
Drive research in efficient machine learning and edge AI for speech and audio applications.
Design compact model architectures capable of operating under strict latency, memory, and power constraints.
Develop and improve state-of-the-art approaches for:
Knowledge distillation
Model pruning
Quantization
Low-rank adaptation and compression
Hardware-aware architectures
Efficient training and inference techniques
Contribute to speech-to-speech, speech recognition, speech translation, text-to-speech, and audio generation systems.
Inference Optimization
Build highly optimized inference pipelines for mobile and embedded hardware.
Improve performance across CPUs, GPUs, NPUs, and other acceleration hardware.
Optimize:
Memory utilization
Operator execution
Kernel performance
Scheduling strategies
Caching mechanisms
End-to-end system latency
Integrate models with production runtimes and deployment frameworks.
Performance Evaluation
Develop rigorous benchmarking methodologies for edge AI systems.
Measure and improve:
Real-time factor
End-to-end latency
Time-to-first-audio
Model size
Peak memory consumption
Power efficiency
Thermal behavior
Speech quality and accuracy
Validate performance directly on target devices rather than relying solely on simulator environments.
Cross-Functional Collaboration
Partner with machine learning researchers, mobile engineers, and systems engineers to bring research into production.
Translate research prototypes into scalable products and customer-facing technologies.
Communicate findings through internal documentation, technical publications, conference papers, and open-source contributions where appropriate.
Required Qualifications
Master's degree, Ph.D., or equivalent industry experience in:
Machine Learning
Speech Processing
Computer Science
Computer Engineering
Efficient Deep Learning
Related quantitative disciplines
Demonstrated expertise in model compression, efficient inference, or edge AI through research publications, production systems, or both.
Strong understanding of one or more of the following:
Quantization
Knowledge distillation
Pruning
Low-rank methods
Efficient neural architectures
Hardware-aware optimization
Strong software engineering skills with:
PyTorch or JAX
C/C++ or equivalent systems-level programming experience
Experience optimizing neural networks for resource-constrained hardware.
Practical knowledge of:
Mobile CPUs
GPUs
NPUs
Memory systems
Numerical precision tradeoffs
Ability to make informed tradeoffs between model quality, latency, memory footprint, power consumption, and deployment portability.
Professional proficiency in English.
Ability to thrive in a fast-moving, research-driven environment.
Preferred Qualifications
Experience working with:
Automatic Speech Recognition (ASR)
Text-to-Speech (TTS)
Speech Translation
Speech-to-Speech Models
Audio-Language Models
Neural Audio Codecs
Familiarity with deployment frameworks such as:
Core ML
ExecuTorch
ONNX Runtime
LiteRT / TensorFlow Lite
TensorRT
llama.cpp
Experience with acceleration technologies including:
Metal
Vulkan
CUDA
QNN
XNNPACK
Custom kernels and operator fusion
Knowledge of:
INT8 quantization
IN .