6 months ago
Base Salary
$140k - $200k/yr
Responsibilities
- Design, build, and maintain sandboxed RL environments including terminal emulators, browser automation harnesses, computer-use simulators, and tool-augmented workspaces.
- Develop reproducible, containerized execution environments using Docker, virtual machines, and lightweight sandboxes for deterministic rollouts and reward collection.
- Integrate and extend open-source agentic tooling and custom CLI/API harnesses for multi-step agent interaction.
- Build structured logging, trajectory capture, state snapshotting, and other observability layers for auditable training and annotation data.
- Collaborate with data operations on task curricula and evaluation protocols across environment types.
- Own environment deployment and reliability through CI/CD pipelines, automated configuration testing, and monitoring for drift or breakage.
- Rapidly prototype new environment types and deliver tested, documented systems from evolving requirements.
Requirements
- At least 2 years of professional software engineering experience.
- Strong Python skills and proficiency in at least one systems-level language such as Go, Rust, or C++.
- Production or near-production experience with containerization and sandboxing tools such as Docker, Podman, or Firecracker.
- Working knowledge of reinforcement learning concepts including MDPs, reward shaping, episode structure, and observation/action spaces.
- Experience building or maintaining developer tooling, CLI tools, or infrastructure automation.
- Comfort with browser automation frameworks or terminal interaction tooling.
- Ability to debug failures across process boundaries, container layers, and network calls.
- Ability to implement from academic papers and open-source benchmark repositories independently.
- Preferred: experience building or contributing to RL environments using Gymnasium, Gym, PettingZoo, or custom implementations.
- Preferred: experience with agentic AI evaluation frameworks such as SWE-bench, WebArena, OSWorld, or TerminalBench.
- Preferred: familiarity with GCP or AWS infrastructure, including Compute Engine, ECS/EKS, or Cloud Build.
- Preferred: prior experience at an AI data company, ML platform company, or AI research lab.
- Preferred: open-source contributions in reinforcement learning, agents, or developer tools.
Benefits
- Hybrid work model with 2 days per week in the office.
- Work from Labelbox tech hubs in San Francisco or Wrocław, Poland.
- Career advancement tied to individual impact.
- Startup-level ownership with growth-stage resources.
- Work across projects for AI labs and model capabilities without single-product monotony.
About Labelbox
Labelbox builds a data platform and managed services for creating, managing, and evaluating training data for AI models, including computer vision, NLP, and RLHF workflows. The company sells subscriptions and labeling services to enterprises and AI labs; founded in 2018 and headquartered in San Francisco, it is privately held and used by Fortune 500 customers.
