Clera

Research Engineer, Synthetic Data

Clera
Apply
3 days ago
Singapore, SingaporeMid Level / Senior

Base Salary

$150k - $250k/yr

Responsibilities

  • Build end-to-end pipelines that generate realistic, structured, and challenging synthetic training tasks.
  • Collaborate with subject-matter experts to create synthetic tasks for AI agents across professional and technical domains.
  • Develop systems to generate, mutate, validate, and continuously improve structured datasets at scale.
  • Analyze model and agent performance and develop metrics for task diversity, realism, learnability, and quality.
  • Design evaluation frameworks, benchmarks, and testing environments for AI agents and large language models.
  • Own and deliver technical projects end-to-end in an early-stage, minimally specified environment.

Requirements

  • 2–4 years of experience in software engineering, machine learning engineering, or AI research.
  • Hands-on experience building data pipelines, ML infrastructure, or synthetic data systems.
  • Proficiency in Python and Linux environments, including containerization with Docker.
  • Experience applying synthetic-data research methods to build end-to-end AI/ML data-generation pipelines.
  • Strong understanding of synthetic-data quality criteria, evaluation metrics, and their limitations.
  • Experience building or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or large language models.
  • Experience automating the generation, validation, mutation, or processing of structured datasets at scale.
  • Familiarity with reinforcement learning training paradigms, agentic AI workflows, or LLM post-training pipelines is a plus.
  • Ability to independently own technical projects, identify edge cases and quality issues, and communicate effectively across time zones.

Benefits

  • Visa sponsorship is available.
  • The role is on-site in Singapore.

Categories

Data Engineering
Contact me