Genesis Molecular AI

Machine Learning Infrastructure Engineer

Genesis Molecular AI
Apply
10 months ago
Remote, United States +2 moreSenior

Responsibilities

  • Design and evolve the in-house workflow orchestration framework, including DAG construction, execution, dependency management, scheduling, caching, retries, observability, and distributed execution.
  • Build and optimize large-scale preprocessing pipelines for protein structures, chemical datasets, simulations, and machine learning training data.
  • Profile end-to-end workflows and eliminate bottlenecks involving I/O, serialization, recomputation, and parallelization.
  • Develop lazy execution, incremental computation, intelligent caching, and artifact-reuse abstractions.
  • Partner with ML researchers, computational chemists, and scientific software engineers to create scalable and reproducible computational pipelines.
  • Improve the developer experience for authoring, debugging, monitoring, and extending distributed scientific workflows.
  • Make architectural decisions involving local and distributed execution, storage, compute scheduling, data lineage, and reproducibility.

Requirements

  • Strong systems engineering skills across APIs, distributed systems, storage, serialization, concurrency, resource scheduling, and performance.
  • Hands-on experience building production pipelines or infrastructure that processes large datasets reliably and efficiently.
  • Ability to work across workflow-framework and application layers and build pragmatic abstractions for specialized scientific workflows.
  • Ability to collaborate closely with researchers and translate evolving scientific workflows into robust infrastructure.
  • Nice to have experience with large-scale ETL, data processing, or workflow systems such as Apache Spark, GCP Dataflow, Apache Beam, Flyte, Dagster, Airflow, or Ray.
  • Nice to have experience designing workflow engines, schedulers, DAG execution systems, build systems, or dependency-driven computation frameworks.
  • Nice to have experience optimizing scientific, ML, protein, cheminformatics, or computational biology data pipelines.
  • Nice to have familiarity with Kubernetes, containerized compute, Terraform, cloud object storage, distributed compute, high-performance or GPU-based computing, and large-scale caching or data-lineage systems.
  • Nice to have familiarity with molecular data formats and tools such as RDKit, OpenEye, BioPython, molecular dynamics tooling, or structural biology pipelines.

Benefits

  • Competitive compensation package including salary and equity.
  • Medical, dental, and vision benefits covered 100% for employees.
  • 401(k) plan.
  • Open unlimited PTO policy.
  • Free lunches and dinners at the offices.
  • Paid maternity and paternity leave.
  • Life and long- and short-term disability insurance.
  • The company is headquartered in San Mateo, California, with an integrated laboratory in San Diego, California.

Tech Stack

Categories

BackendData Engineering
Genesis Molecular AI

About Genesis Molecular AI

51-200 employees
Contact me