8 months ago
Responsibilities
- Build and scale audio-model training data pipelines, including preprocessing, augmentation, and quality filtering
- Design, implement, and maintain evaluation systems for multimodal performance across internal and public benchmarks
- Fine-tune and adapt audio models for customer-specific use cases from requirements through deployment
- Contribute production code to the core audio repository while collaborating with infrastructure and research teams
- Support experimentation under real hardware constraints and shift between customer work and core development as priorities change
- Own customer workstreams and deliver production-ready pipelines or evaluation systems
Requirements
- Strong programming fundamentals and the ability to write clean, maintainable, production-grade code
- Experience building and shipping production ML systems beyond model training, including data pipelines, evaluation systems, or serving infrastructure
- Proficiency in PyTorch and familiarity with distributed training frameworks such as DeepSpeed or FSDP
- Experience collaborating in shared codebases with high engineering standards
- Direct experience with audio or speech models such as ASR, TTS, vocoders, diarization, or speech-to-speech systems is preferred
- Experience running large-scale training experiments on distributed GPU clusters is preferred
- Open-source contributions demonstrating code quality and engineering judgment are preferred
Benefits
- 100% of medical, dental, and vision premiums covered for employees and dependents
- 401(k) matching up to 4% of base pay
- Unlimited PTO and company-wide Refill Days
- Equity offered
- Work involves production systems spanning data-center accelerators and on-device hardware
Tech Stack
Categories
Data EngineeringML Engineering
About Liquid AI
We build efficient general-purpose AI at every scale.
