17 hours ago
Singapore, SingaporeMid Level / Senior
Base Salary
$150k - $250k/yr
Responsibilities
- Design and build end-to-end synthetic data pipelines that turn domain-specific workflows into realistic training tasks.
- Collaborate with subject-matter experts to create synthetic tasks for AI agents across professional and technical domains.
- Develop methods to maximize synthetic task diversity, realism, and learnability.
- Build tooling to mutate, validate, and iteratively improve synthetic tasks at scale.
- Analyze model and agent performance to identify what synthetic tasks teach and where agents break down.
- Define and implement metrics for synthetic task quality across diversity, realism, and learnability.
Requirements
- Require 2 to 4 years of experience in software engineering, machine learning engineering, or AI research, focused on data pipelines, ML infrastructure, or synthetic data systems.
- Require proficiency in Python and hands-on experience with Docker and Linux environments.
- Require demonstrated experience applying synthetic data research methods to build generation pipelines end-to-end.
- Require understanding of synthetic data quality criteria, including diversity, realism, learnability, and their limitations.
- Require experience designing, implementing, or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or large language models.
- Require a track record of independently owning and delivering technical projects with minimal predefined requirements or roadmap.
- Require strong attention to edge cases, inconsistencies, and quality issues in synthetic or algorithmically generated datasets.
- Familiarity with reinforcement learning paradigms, agentic AI workflows, or LLM post-training pipelines is a plus.
- Strong communication skills for collaboration across time zones are required.
Benefits
- Salary range is $150,000 to $250,000 USD annually.
- Visa sponsorship is available.
- The role is on-site in Singapore.
Categories
Data EngineeringML Engineering
