
Staff Engineer, Data Platform
Lila Sciences3 months ago
Base Salary
$192k - $272k/yr
Responsibilities
- Set the technical direction and architecture for core data infrastructure.
- Design and evolve systems that ingest, store, transform, and serve data across scientific and ML workflows.
- Build reliable ingestion pipelines for laboratory instruments, public scientific datasets, and external research literature.
- Own interfaces between upstream data producers and downstream consumers.
- Operate and extend workflow orchestration systems for complex scientific pipelines.
- Establish observability, fault tolerance, and reproducibility across the data stack.
- Define data models, schema evolution practices, and data contracts.
- Partner with ML researchers, lab scientists, and product engineers to translate requirements into platform capabilities.
- Establish coding, review, and design standards for the data platform team.
- Mentor engineers, lead design reviews, and raise the technical bar across the group.
Requirements
- Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
- 8+ years of experience as a software or data engineer focused on building and operating data infrastructure.
- Experience designing and shipping ingestion frameworks, storage abstractions, and orchestration systems from the ground up.
- Fluency in Python and SQL with the ability to write production-quality code.
- Production experience with relational and NoSQL databases, schema design, query optimization, and operating systems at scale.
- Experience working with structured, semi-structured, and unstructured data.
- Experience with cloud infrastructure and containerized deployment using AWS and Kubernetes.
- Hands-on experience with modern table formats and open lakehouse patterns, including Iceberg, Delta Lake, and Hudi.
- Strong cross-functional experience translating requirements from scientists, ML researchers, and engineers into platform decisions.
- Bonus: experience with Flyte, Airflow, Dagster, or similar workflow orchestration systems.
- Bonus: experience building data infrastructure for agentic and LLM-driven workflows, including vector databases, RAG infrastructure, and retrieval-optimized data access patterns.
- Bonus: background in scientific computing, life sciences, or research software.
- Bonus: proficiency with AI-assisted development tools such as Cursor or Claude Code.
Benefits
- Medical, dental, and vision coverage for full-time U.S. employees
- Employer-paid life and disability insurance
- Flexible time off and generous company-wide holidays
- Paid parental leave
- Educational assistance program
- Commuter benefits, including bike share memberships for office-based employees
- Company-subsidized lunch program
- Regionally tailored benefits for full-time employees outside the U.S.
Tech Stack
Categories
BackendData Engineering
About Lila Sciences
Lila Sciences is the world’s first scientific superintelligence platform and autonomous lab for life, chemistry, and materials science. We are building the foundation to apply AI to every aspect of the scientific method, enabling scientists to bring forth solutions in human health and sustainability at a pace and scale never experienced before.