
AI/ML Research Engineer, LLM Post-Training & Evaluation
Innodata Inc.1 month ago
Remote, CanadaMid Level
Responsibilities
- Lead or co-lead technically complex ML engineering projects from customer discussions through implementation and delivery.
- Design and improve LLM training, post-training, data ingestion, preprocessing, fine-tuning, evaluation, and experiment-tracking pipelines.
- Implement evaluation systems, offline benchmarks, and task-specific test harnesses for LLMs and multimodal models.
- Integrate human-in-the-loop and AI-augmented evaluation signals into model development workflows.
- Build infrastructure for reproducible experimentation, metrics logging, regression monitoring, and scalable model assessment.
- Diagnose model behavior and pipeline failures, including data issues, training instability, metric inconsistencies, and evaluation drift.
- Collaborate with research scientists, language data scientists, data engineers, and customer technical stakeholders.
- Contribute to internal research, benchmark datasets, evaluation tooling, post-training workflows, documentation, design reviews, and engineering standards.
- Mentor junior engineers and explain complex technical tradeoffs to technical and non-technical audiences.
Requirements
- Bachelor’s, master’s, or doctoral degree in Computer Science, Machine Learning, AI, Applied Mathematics, or a related quantitative technical field; MS or PhD preferred.
- 2–3 years of relevant industry or research engineering experience in ML/AI systems.
- Hands-on experience with LLM training, fine-tuning, or post-training, including supervised fine-tuning, preference optimization, RLHF/RLAIF workflows, or foundation-model adaptation.
- Strong Python programming and production-quality ML software engineering skills.
- Experience with PyTorch, JAX, TensorFlow, the Hugging Face ecosystem, vLLM, or distributed training tooling.
- Experience designing evaluation pipelines, metrics computation, dataset handling, experiment comparison, automated test harnesses, and reproducibility practices.
- Understanding of data pipelines, ML systems engineering, observability, debugging, scalable data processing, and workflow orchestration.
- Experience with distributed ML systems and performance optimization for training or evaluation workloads, preferably in GPU or accelerator environments.
- Understanding of transformer-based models, post-training workflows, inference latency, throughput, and memory/performance tradeoffs.
- Ability to collaborate with research scientists, ML engineers, data engineers, and customer technical leads.