about 3 hours ago
New York, NY, USAMid Level
Base Salary
$175k - $275k/yr
Responsibilities
- Train and optimize large-scale video and multimodal models.
- Improve training and inference efficiency across memory, latency, and cost.
- Implement distillation, quantization, and pruning techniques to accelerate diffusion and autoregressive generation.
- Build and maintain distributed training systems.
- Optimize GPU utilization, parallelism, and throughput.
- Develop tooling for experimentation, evaluation, and debugging.
- Translate research models into robust, production-ready systems.
- Monitor and improve model performance in real-world usage.
Requirements
- BS, MS, or PhD in computer science, machine learning, or a related field.
- At least two years of professional industry experience.
- Strong experience with deep learning systems and infrastructure.
- Expertise in PyTorch, CUDA, Triton, and distributed training, including FSDP.
- Experience scaling and optimizing large models under low-latency inference constraints.
- Strong debugging and performance-profiling skills.
- Ability to move quickly from prototype to production.
Benefits
- Comprehensive medical, dental, and vision plans; 401(k) with employer match; commuter benefits; catered lunch multiple days per week; nightly dinner stipend when working late; Grubhub subscription; health and wellness perks; team offsites and monthly team events; generous PTO.
- The role is full-time and requires working in person at Mirage’s NYC headquarters in Union Square.
About Mirage
At Mirage, we’re building full-stack foundation models and products that redefine video creation. Over 20 million content creators and businesses use Captions to reach their full creative and commercial potential.
