
AI/ML Research Engineer, LLM Post-Training & Evaluation
Innodata Inc.2 months ago
Remote, United StatesMid Level
Base Salary
$80k - $175k/yr
Responsibilities
- Lead or co-lead technically complex ML engineering projects from customer discussions through implementation and delivery.
- Design, build, and improve LLM training, fine-tuning, post-training, evaluation, data ingestion, preprocessing, and experiment-tracking pipelines.
- Implement evaluation systems for LLMs and multimodal models, including offline benchmarks and task-specific test harnesses.
- Integrate human-in-the-loop and AI-augmented evaluation signals into model development workflows.
- Build infrastructure and tooling for reproducible experimentation, metrics logging, and regression monitoring.
- Diagnose data issues, training instability, metric inconsistencies, evaluation drift, model behavior, and pipeline failures.
- Collaborate with language data scientists, applied research scientists, data engineers, and customer technical stakeholders.
- Contribute to benchmark datasets, evaluation frameworks, post-training workflows, platform development, documentation, technical design reviews, and engineering standards.
- Mentor junior engineers and explain technical tradeoffs to technical and non-technical audiences.
Requirements
- Bachelor’s, master’s, or PhD in Computer Science, Machine Learning, AI, Applied Mathematics, or a related quantitative technical field; MS or PhD preferred.
- 2–3 years of relevant industry or research engineering experience in ML or AI systems.
- Hands-on experience with LLM training, fine-tuning, or post-training, including supervised fine-tuning, preference optimization, RLHF/RLAIF workflows, or foundation-model adaptation.
- Strong Python programming skills and production-quality software engineering fundamentals.
- Experience with PyTorch, JAX, TensorFlow, the Hugging Face ecosystem, vLLM, or distributed training stacks.
- Experience designing automated LLM/ML evaluation pipelines, metrics computation, dataset handling, experiment comparisons, and test harnesses.
- Understanding of reproducibility, observability, debugging, versioning, and experiment tracking for ML systems.
- Experience with distributed ML systems, performance optimization, large-scale data processing, and workflow orchestration.
- Familiarity with inference latency, throughput, memory and performance tradeoffs, data-processing pipelines, storage formats, scalable datasets, CI/CD, testing, and ML engineering quality practices.
- Ability to collaborate with research scientists, ML engineers, data engineers, and customer technical leads.
Categories
AI ResearchML Engineering