10 days ago
San Diego, CA, USASenior / Staff+
H1B sponsor
Base Salary
$60k - $149k/yr
Responsibilities
- Build and integrate evaluation harnesses and automation for software development use cases, including converting merged pull requests and other engineering artifacts into repeatable benchmark tasks.
- Develop versioned and reproducible evaluation workflows with pinned dependencies, containerized runs, isolated worktrees, automated test execution, and result validation.
- Validate and calibrate evaluation methods against human judgment to ensure scores are both consistent and accurate.
- Execute benchmarks across quality, productivity, efficiency, cost, and latency measures.
- Analyze repeated-run variance, failure patterns, cost per outcome, and workflow reliability while improving automation.
- Collaborate with engineering and data teams and document evaluation methods and findings for technical and leadership audiences.
Requirements
- Strong software engineering background with experience building automation, developer tooling, or test and validation systems.
- Experience evaluating AI coding agents that modify code repositories and validating generated code changes against expected outcomes.
- Experience building automated, reproducible evaluation harnesses and benchmarking workflows.
- Experience establishing baselines, measuring run-to-run variance, analyzing failures, and improving evaluation reliability.
- Experience with Git repository history, branches, pull requests, and automated testing in CI/CD pipelines.
- Experience using LLMs as judges for coding-agent outputs and calibrating results against human assessments.
- Proficiency in at least one general-purpose language such as Python, Java, or JavaScript.
- Working knowledge of Git and Docker, including branches, history, working trees, and containerization.
- Experience with APIs, development environments, CI/CD pipelines, and standard engineering workflows.
- Understanding of AI, LLM, and agent evaluation issues, including judge validity, misleading single runs, and benchmark contamination.
- Hands-on experience with AI coding tools and agentic harnesses such as Claude Code, Devin, or Cursor.
- Five to eight years of experience.
- Mandatory skill: Cloud Product & Platform Testing.
Benefits
- Wipro’s standard benefits include medical and dental options, disability insurance, paid time off including sick leave, and other paid and unpaid leave options.
- The expected base compensation ranges from $60,000 to $148,500, with final compensation depending on location, minimum wage obligations, skills, and relevant experience.
- Some roles may require successful completion of a post-offer drug screening where permitted by applicable state law.
Tech Stack
Categories
About Wipro
Wipro is a global IT services and consulting firm that provides application development, cloud migration, infrastructure management, data/AI, and business process outsourcing for enterprises. It earns revenue through long-term outsourcing, managed services, and project-based consulting across multiple industries. Founded in 1945 and headquartered in Bangalore, India, Wipro is publicly traded (NYSE: WIT; NSE/BSE: WIPRO).
