6 hours ago
Base Salary
$500k - $850k/yr
Responsibilities
- Design and implement highly performant, reproducible, and traceable data-processing infrastructure for large language model training.
- Develop scalable processing primitives including tokenization, deduplication, and chunking.
- Build data quality assurance and validation systems at scale.
- Collaborate with research teams to implement novel data-processing architectures.
- Build and operate end-to-end pipelines that transform raw web-scale corpora into training-ready datasets.
- Develop distributed computing architectures and fault-tolerant infrastructure components based on research requirements.
Requirements
- At least 5 years of professional experience outside internships.
- Strong software engineering skills building high-throughput, fault-tolerant distributed systems.
- Hands-on experience with distributed computing frameworks, particularly Apache Spark.
- Advanced degree in Computer Science or a related field preferred.
- Experience with language model training infrastructure preferred.
- Background in data infrastructure, MLOps, or ML infrastructure preferred.
- Expertise with Python and Rust preferred.
- Strong problem-solving, communication, collaboration, ownership, and attention-to-detail skills.
Benefits
- Hybrid work policy requiring staff to be in an Anthropic office at least 25% of the time, with some roles requiring more office time.
- Visa sponsorship with immigration-lawyer support when sponsorship is possible.
- Competitive compensation and benefits.
- Optional equity donation matching.
- Generous vacation and parental leave.
- Flexible working hours.
- Collaborative office space.
Tech Stack
Categories
Data EngineeringML Engineering
About Anthropic
We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.
