
Principal ML Ops Engineer (EMEA Remote)
Pragmatike2 months ago
Prague, Czechia +7 moreStaff+
Responsibilities
- Build and operate production-grade model-serving infrastructure using vLLM, TGI, Triton, or equivalent frameworks.
- Design deployment pipelines with blue/green and canary rollout strategies for ML models.
- Develop autoscaling systems, multi-model serving architectures, and intelligent request-routing layers.
- Optimize GPU utilization, memory efficiency, network throughput, and model artifact storage performance.
- Design observability for inference latency, throughput, GPU usage, cost metrics, and system health.
- Manage model registries and automated, reproducible model deployment pipelines.
- Own the ML systems lifecycle from development through production, including operational support and on-call responsibilities.
- Define engineering best practices and contribute to platform scalability.
Requirements
- At least four years of experience in MLOps, platform engineering, SRE, or similar infrastructure roles focused on ML systems.
- Hands-on experience with model-serving frameworks such as vLLM, TGI, Triton, or equivalent.
- Strong experience with container orchestration and production GPU-based workloads.
- Experience with MLOps tooling, including model registries, experiment tracking, and automated deployment pipelines.
- Proficiency in Python and infrastructure-as-code tools such as Terraform and Helm.
- Strong understanding of distributed systems, performance tuning, and production reliability engineering.
- Ability to use AI coding assistants for development and debugging workflows.
- Preferred experience with Kubeflow, MLflow, or KubeAI.
- Preferred knowledge of GPU scheduling, CUDA or ROCm optimization, and multi-tenant inference systems.
- Preferred experience with GPU cost optimization, early-stage startups, greenfield infrastructure, and production systems built from scratch.
- Fluent English is required.
Benefits
- Fully remote role for candidates working in EMEA time zones
- ASAP start date
- Opportunity to build foundational ML inference infrastructure from the ground up
- Work at the intersection of distributed systems, GPU computing, and sustainable cloud architecture
- Influence core engineering decisions and define scalable engineering best practices
Categories
About Pragmatike
Trusted by remote-first companies worldwide. Completing tech projects for startups and scaleups.