
Software Engineer, Data Infrastructure
Thinking Machines Lab5 days ago
Base Salary
$350k - $475k/yr
Responsibilities
- Design, build, and operate scalable, fault-tolerant infrastructure for distributed compute, data orchestration, and multimodal storage.
- Develop high-throughput systems for data ingestion, processing, transformation, training data catalogs, deduplication, quality checks, and search.
- Build traceability, reproducibility, and quality-control systems across the data lifecycle.
- Implement and maintain monitoring and alerting for platform reliability and performance.
- Collaborate with research teams to improve data quality, enable new features, and accelerate training cycles.
Requirements
- Bachelor’s degree or equivalent experience in computer science, engineering, or a similar field.
- Proficiency in at least one backend language, specifically Python or Rust.
- Fluency with distributed compute frameworks such as Apache Spark or Ray.
- Deep familiarity with cloud infrastructure, data lake architectures, and batch and streaming pipelines.
- Ability to operate across the stack and own projects end-to-end.
- Preferred experience with Kafka, dbt, Terraform, and Airflow.
- Preferred experience building a web crawler and scaling deduplication, data mining, and search systems.
- Strong knowledge of file formats and storage systems such as Parquet and Delta Lake, including their performance and scalability implications.
- Proactivity around documentation, testing, and building tooling that empowers teammates.
Benefits
- The role is based in San Francisco, California.
- Visa sponsorship is available.
- Benefits include health, dental, and vision coverage, unlimited PTO, paid parental leave, and relocation support as needed.
Tech Stack
Categories
BackendData Engineering