Member of Technical Staff, LLM Evaluation Infra
The Inception Company6 months ago
San Mateo, CA, USAMid Level
Responsibilities
- Build scalable automated evaluation pipelines integrated with model training and deployment workflows.
- Design and maintain evaluation frameworks and benchmarks for LLM performance across diverse tasks and domains.
- Conduct statistical analysis of model outputs to identify failure modes, biases, and performance gaps.
- Partner with product and customer-facing teams to translate real-world use cases into evaluation criteria.
- Define quantitative metrics for model quality, safety, reliability, and regression detection.
Requirements
- Bachelor's, master's, or doctoral degree in Computer Science, Machine Learning, Statistics, or a related field, or equivalent experience.
- At least 2 years of experience in ML evaluation, applied ML research, or a related engineering role.
- Experience with Git, Docker, and cloud services such as AWS, GCP, or Azure.
- Understanding of LLM fundamentals including autoregressive generation, instruction tuning, RLHF, in-context learning, and decoding strategies.
- Proficiency in Python and ML frameworks such as PyTorch.
- Experience designing evaluation metrics and benchmarks for generative models.
- Strong communication skills for translating complex evaluation results into actionable insights.
- Preferred qualifications include familiarity with benchmark suites, agentic evaluations such as SWE-bench and GPQA/GDPval, Kubernetes, Terraform, MLOps, data engineering, large-scale data labeling, synthetic data generation, LLM safety and alignment evaluation, and human-in-the-loop evaluation systems.