about 3 hours ago
Base Salary
$180k - $240k/yr
Responsibilities
- Define and build evaluation methodologies for various AI models.
- Design, build, and maintain automated evaluation pipelines for batch and streaming environments.
- Create scalable evaluation infrastructure running against production models.
- Translate research benchmarks into automated pass/fail criteria.
- Build canaries and monitoring systems to detect quality regressions.
- Collaborate with DevOps to establish test environments and infrastructure.
- Integrate evaluation gates into CI/CD processes for continuous quality verification.
- Participate in code reviews and technical discussions to enhance engineering practices.
Requirements
- BS, MS, or PhD in Computer Science, AI, Applied Math, or related field.
- 5+ years of professional software or QA engineering experience.
- Solid backend/scripting experience in languages like Python, Rust, or Go.
- Experience in designing automated test pipelines or evaluation frameworks.
- Strong analytical skills to assess metrics and statistical variations.
- Ability to tackle ambiguous technical challenges and communicate effectively.
