Anthropic

Full-Stack Software Engineer, Reinforcement Learning

Anthropic
Apply
5 months ago
San Francisco, CA, USA or New York, NY, USAMid Level
H1B Sponsor

Base Salary

$1 - $2/yr

Responsibilities

  • Build web platforms for reinforcement-learning environment creation, management, versioning, validation, and quality review.
  • Develop vendor-facing tools for creating, submitting, and iterating on training environments.
  • Build human data-collection platforms with labeling workflows, quality assurance, and feedback mechanisms.
  • Create evaluation dashboards and observability UIs for environment quality, training health, and reward hacking.
  • Develop backend services and APIs connecting environment authoring, data collection, and RL training infrastructure.
  • Build scalable code-data generation pipelines producing programming tasks with robust reward signals.
  • Create onboarding automation and documentation tooling for vendors and internal users.
  • Partner with RL researchers, data operations, and vendor management to turn ambiguous needs into well-designed products.

Requirements

  • Strong software engineering fundamentals and full-stack experience, including ownership from database schema through frontend.
  • Proficiency in Python and a modern web stack such as React and TypeScript.
  • Demonstrated ability to ship systems that solve difficult problems and materially improve team effectiveness.
  • Ability to work independently, handle ambiguous requirements, and drive projects forward.
  • Strong UX judgment for interfaces used by technical researchers and non-technical labelers.
  • Clear communication with researchers, operations teams, engineers, and external partners.
  • Preferred experience with data collection, labeling, or annotation platforms at scale.
  • Preferred experience with multi-tenant platforms, role-based access, audit trails, and vendor management workflows.
  • Preferred experience with GCP or AWS, Docker, and CI/CD pipelines.
  • Preferred familiarity with LLM training, fine-tuning, or evaluation workflows.
  • Preferred experience with async Python using Trio or asyncio, high-throughput API design, dashboards, monitoring, or observability tooling.

Benefits

  • Hybrid policy requiring staff to be in an office at least 25% of the time.
  • Visa sponsorship may be available, with immigration lawyer support.
  • Competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration space.
Anthropic

About Anthropic

501-1,000 employees

We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.

Contact me