Base Salary
$180k - $220k/yr
Responsibilities
- Build and own an AI-powered ingestion and normalization pipeline for Excel, CSV, laboratory, instrument, and internal pipeline data.
- Develop schema mapping, coercion, unit normalization, metadata standardization, variable-name harmonization, batch correction, and vendor-format conversion logic.
- Use LLM-driven and classical data-engineering tools to extract metadata, infer column roles and types, clean headers, resolve inconsistencies, and prepare canonical datasets.
- Ensure one-time transformations occur during ingestion so downstream analytics and AI systems receive clean data.
- Build validation, verification, and quality-control layers to identify ambiguous, inconsistent, or corrupt data.
- Collaborate with product, data science, bioinformatics, and infrastructure teams to define data standards and integrate pipeline outputs with analytics and storage systems.
Requirements
- 5+ years of experience in data engineering or data wrangling with real-world tabular or semi-structured data.
- Strong Python fluency and experience with data-processing tools such as Pandas, Polars, or PyArrow.
- Extensive experience handling and normalizing messy Excel, CSV, and spreadsheet-style data.
- Experience designing and maintaining robust ETL/ELT pipelines, ideally for scientific or laboratory data.
- Ability to combine classical data engineering with LLM-powered data normalization, metadata extraction, or cleaning.
- Ability to own ingestion and normalization end-to-end with a focus on maintainability, reproducibility, and scalability.
- Strong communication and cross-functional collaboration skills.
- Familiarity with scientific data types such as plate-reader data, genomics metadata, time-series data, batch information, and instrumentation outputs is preferred.
- Experience with workflow orchestration tools such as Nextflow, Prefect, Airflow, or Dagster is preferred.
- Experience with cloud infrastructure, AWS S3, data lakes or warehouses, and database schemas is preferred.
- Experience building or integrating LLM-based data transformation or cleansing agents is preferred.
- Computational biology, laboratory data, or bioinformatics experience is a bonus but not required.
Benefits
- Mission-driven opportunity to improve the quality and reliability of scientific data and AI systems.
- High ownership and autonomy over ingestion architecture, standards, and pipelines.
- Collaboration with a tight-knit team of engineers, scientists, and builders.
- In-person role based in a San Francisco office.
- Comprehensive PPO medical, dental, and vision coverage through Anthem.
- 401(k) with top-tier plans.
Tech Stack
Categories
About Mithrl
The scientific decision engine for R&D. Raw data → defensible decisions, faster. Pharma and biotech R&D often spend weeks, or even months, learning, coding, and troubleshooting discovery analysis pipelines. This slows research and diverts focus from the work that truly matters: discovery. Mithrl changes that. Using only natural language from the user, Mithrl AI builds custom workflows with full transparency and reproducibility for discovery data (omics and beyond), on-demand and in minutes, not weeks. Discovery researchers across therapeutic areas can instantly align, analyze, interpret, integrate, and iterate hypotheses for complex datasets without manual coding or repetitive setup. By automating these time-consuming processes, Mithrl empowers scientists to spend more time designing experiments, testing insights, and driving high-quality discoveries, including novel IP. From accelerating data analysis to uncovering novel targets and biomarkers, Mithrl is the Scientific Decision Engine that turns data into actionable breakthroughs. See Mithrl in action. Book a demo and transform your research today! Website: www.mithrl.com Follow us on social media: • LinkedIn: linkedin.com/company/mithrl-ai • X: x.com/mithrl_ai
