4 months ago
Stockholm, Sweden or London, United KingdomSenior
Responsibilities
- Own the full lifecycle of Lovable’s post-training pipeline from data curation and training runs through evaluation and production deployment.
- Apply reinforcement learning, preference optimization, and supervised fine-tuning to improve code generation, reasoning about user intent, and reliable agent behavior.
- Build evaluation and experimentation infrastructure covering helpfulness, safety, latency, reliability, and user satisfaction.
- Develop and operate production systems for large-scale training jobs, including GPU orchestration and data pipelines.
- Collaborate with agent, product, and infrastructure engineers to turn model improvements into product improvements.
- Investigate and resolve failures across training recipes, data, inference, serving, and production behavior.
- Read research papers, run experiments, and rapidly move promising research into production.
Requirements
- Personally run post-training jobs on large language models using methods such as RFT/RLVR or preference optimization.
- Write reliable production code and build systems that operate beyond research prototypes.
- Be fluent in at least one major machine learning framework, such as PyTorch or JAX.
- Have experience with distributed training setups and GPU clusters.
- Understand the mathematics of preference optimization, reward modeling, and alignment techniques.
- Have built or substantially contributed to evaluation systems measuring real-world model quality.
- Be able to trace model-quality regressions from user-facing symptoms through serving, inference, and training.
- Preferred qualifications include experience with code generation or agentic workloads, deploying post-trained models to real users at scale, owning data-to-deployment workflows, speculative decoding, predictive evaluation methodology, open-source ML contributions, or relevant publication experience.
Tech Stack
Categories
AI ResearchML Engineering