4 months ago
Base Salary
$140k - $200k/yr
Responsibilities
- Design, build, and maintain sandboxed RL environments including terminal emulators, browser automation harnesses, computer-use simulators, and tool-augmented workspaces.
- Develop reproducible, containerized execution environments using Docker, virtual machines, and lightweight sandboxes for deterministic rollouts and reward collection.
- Integrate and extend open-source agentic tooling and custom CLI/API harnesses for multi-step agent interaction.
- Build structured logging, trajectory capture, state snapshotting, and other observability layers for auditable training and annotation data.
- Collaborate with data operations on task curricula and evaluation protocols across environment types.
- Own environment deployment and reliability through CI/CD pipelines, automated configuration testing, and monitoring for drift or breakage.
- Rapidly prototype new environment types and deliver tested, documented systems from evolving requirements.
Requirements
- At least 2 years of professional software engineering experience.
- Strong Python skills and proficiency in at least one systems-level language such as Go, Rust, or C++.
- Production or near-production experience with containerization and sandboxing tools such as Docker, Podman, or Firecracker.
- Working knowledge of reinforcement learning concepts including MDPs, reward shaping, episode structure, and observation/action spaces.
- Experience building or maintaining developer tooling, CLI tools, or infrastructure automation.
- Comfort with browser automation frameworks or terminal interaction tooling.
- Ability to debug failures across process boundaries, container layers, and network calls.
- Ability to implement from academic papers and open-source benchmark repositories independently.
- Preferred: experience building or contributing to RL environments using Gymnasium, Gym, PettingZoo, or custom implementations.
- Preferred: experience with agentic AI evaluation frameworks such as SWE-bench, WebArena, OSWorld, or TerminalBench.
- Preferred: familiarity with GCP or AWS infrastructure, including Compute Engine, ECS/EKS, or Cloud Build.
- Preferred: prior experience at an AI data company, ML platform company, or AI research lab.
- Preferred: open-source contributions in reinforcement learning, agents, or developer tools.
Benefits
- Hybrid work model with 2 days per week in the office.
- Work from Labelbox tech hubs in San Francisco or Wrocław, Poland.
- Career advancement tied to individual impact.
- Startup-level ownership with growth-stage resources.
- Work across projects for AI labs and model capabilities without single-product monotony.
About Labelbox
Labelbox builds and operates reinforcement learning data factories for the world’s leading AI labs and enterprises, powering the next generation of frontier models and AI applications.
