
Principal Machine Learning Engineer
Digital Turbine, Inc.16 hours ago
Berlin, Germany or New York, NY, USAStaff+
Responsibilities
- Define the long-term enterprise architecture, technical direction, and multi-year roadmap for ML platforms, model deployment infrastructure, and generative AI systems.
- Solve novel engineering challenges involving model scaling, distributed systems, real-time inference, and feature engineering.
- Establish company-wide MLOps practices, security standards, evaluation frameworks, model governance protocols, and operational reliability metrics.
- Architect distributed training clusters, feature platforms, high-throughput model-serving engines, and cost-efficient inference pipelines.
- Lead the technical design and integration of foundation models, LLM orchestration, retrieval-augmented generation, parameter-efficient fine-tuning, and vector infrastructure.
- Design observability systems for model drift, system health, data quality, security posture, and business impact.
- Lead architecture review boards and approve critical system designs, data pipelines, and production deployment architectures.
- Mentor and technically sponsor senior and staff engineers while influencing teams without direct administrative authority.
- Partner with executives, product leaders, ML Engineering, Data Engineering, Infrastructure/DevOps, Enterprise Security, and Product teams on strategic roadmaps and architectural decisions.
Requirements
- Bachelor’s degree in Computer Science, Machine Learning, Data Science, or a related quantitative field; a Master’s or Ph.D. is preferred.
- 10+ years of software or ML engineering experience with a bachelor’s degree, or 8+ years with a Master’s or Ph.D.
- Proven experience operating as a Principal or Staff engineer on enterprise-scale systems.
- Experience architecting, deploying, and maintaining high-throughput, low-latency, mission-critical ML models and pipelines in real-time production environments.
- Expert mastery of PyTorch, TensorFlow, Jax, and modern distributed ML training frameworks such as DeepSpeed, Megatron-LM, and Ray.
- Comprehensive experience with cloud-native infrastructure, Kubernetes, feature stores, CI/CD pipelines, and enterprise MLOps suites.
- Exceptional proficiency in system architecture, distributed computing, memory optimization, data structures, and languages such as Python, C++, or Rust.
- Ability to drive strategic alignment, architectural consensus, and engineering compliance across non-reporting teams and business units.
- Preferred experience includes generative AI architecture, foundation model pre-training and fine-tuning, guardrailing, agentic workflows, vector retrieval, Ray, Apache Spark, Dask, vector databases, NoSQL, distributed caching, GPU clusters, TPUs, custom silicon, CUDA, TensorRT, and ONNX Runtime.
- Preferred industry leadership includes open-source contributions, conference speaking, or peer-reviewed publications in machine learning or distributed systems.
Benefits
- Hybrid work environment; only candidates local to the posting location will be considered.
- Digital Turbine is an equal opportunity employer committed to diversity, inclusion, and an equitable workplace.
Tech Stack
Categories
About Digital Turbine, Inc.
Digital Turbine builds a mobile growth and monetization platform that connects advertisers, app developers, carriers, and OEMs through on-device app discovery, ad delivery, and data-driven optimization. Its products include on-device placements (Ignite), a demand-side platform, offerwall, and content media services that drive user acquisition and ad revenue. The company is publicly traded on NASDAQ (APPS) and is headquartered in Austin, Texas.