24 hours ago
Hyderābād, IndiaSenior
Responsibilities
- Design, develop, validate, and operate foundation-model applications for biomedical and clinical data.
- Adapt and evaluate biological foundation models, protein and sequence models, multimodal models, and large language models for biomarker discovery, target identification, patient stratification, and scientific decision support.
- Build workflows covering data curation, representation learning, fine-tuning or parameter-efficient adaptation, retrieval augmentation, evaluation, and monitored deployment.
- Engineer reproducible pipelines for multi-omics and clinical data with provenance, versioning, quality controls, and access controls.
- Develop agentic workflows combining foundation models, bioinformatics tools, structured knowledge, and human review.
- Define benchmarks and validation strategies addressing biological relevance, robustness, bias, uncertainty, hallucination risk, and reproducibility.
- Develop production-ready services and interfaces using cloud and GPU infrastructure, optimizing performance, cost, reliability, and observability.
- Collaborate with computational biology, wet-lab, clinical, data engineering, and product teams to deliver documented technical solutions.
- Produce technical documentation, model cards, evaluation reports, and methods descriptions for internal review, regulated development, and publication.
- Troubleshoot platform and pipeline issues and promote engineering best practices and responsible AI use.
Requirements
- Master’s or PhD in Bioinformatics, Computational Biology, Computer Science, Machine Learning, Statistics, Genetics/Genomics, or a related discipline.
- At least 7 years of hands-on experience building bioinformatics, machine learning, data science, or research software solutions.
- Strong Python programming and practical software engineering experience, including Git, testing, code review, CI/CD, and documentation.
- Hands-on expertise with deep learning and foundation models, including transformers, self-supervised learning, embedding models, fine-tuning or parameter-efficient adaptation, evaluation, and inference optimization.
- Experience using or adapting biological foundation models for sequence, protein, cellular, molecular, or multimodal biomedical data.
- Familiarity with large language models, retrieval-augmented generation, and tool-using agents.
- Experience with Hugging Face and AWS SageMaker.
- Understanding of genomics, transcriptomics, single-cell or spatial omics, proteomics, imaging, or other biomedical data modalities.
- Experience designing reproducible workflows with Nextflow or Snakemake and containers such as Docker or Singularity.
- Experience with cloud and HPC environments, GPU compute, distributed training or inference, and scalable data processing frameworks.
- Working knowledge of FASTQ, BAM/CRAM, VCF/MAF, HDF5, AnnData, Seurat, and biological metadata practices.
- Experience curating, integrating, and governing data from resources including TCGA, GTEx, GEO, SRA, dbGaP, cBioPortal, ClinVar, CellxGene, COSMIC, gnomAD, and UniProt.
- Ability to design meaningful benchmarks and communicate model performance, limitations, uncertainty, and responsible-use guidance.
- Strong statistical reasoning and experience with quality control and evaluation methods for biological data and machine-learning systems.
- Biomedical, pharmaceutical, or regulated research experience is preferred.
Benefits
- Full-time employment at the Amgen India office in Hyderabad.
Categories
About Amgen
Amgen is a public biotechnology company that discovers, develops, manufactures, and sells biologic and biosimilar therapeutics for patients with serious diseases. Its portfolio and pipeline span oncology, cardiovascular disease, osteoporosis, inflammation, and rare diseases, marketed globally through direct sales and partnerships. Founded in 1980 and headquartered in Thousand Oaks, California (NASDAQ: AMGN), Amgen is a component of the Dow Jones Industrial Average and completed the acquisition of Horizon Therapeutics in 2023.
