Pragmatike

Principal ML Ops Engineer (EMEA Remote)

Pragmatike
Apply
2 months ago
Prague, Czechia +7 moreStaff+

Responsibilities

  • Build and operate production-grade model-serving infrastructure using vLLM, TGI, Triton, or equivalent frameworks.
  • Design deployment pipelines with blue/green and canary rollout strategies for ML models.
  • Develop autoscaling systems, multi-model serving architectures, and intelligent request-routing layers.
  • Optimize GPU utilization, memory efficiency, network throughput, and model artifact storage performance.
  • Design observability for inference latency, throughput, GPU usage, cost metrics, and system health.
  • Manage model registries and automated, reproducible model deployment pipelines.
  • Own the ML systems lifecycle from development through production, including operational support and on-call responsibilities.
  • Define engineering best practices and contribute to platform scalability.

Requirements

  • At least four years of experience in MLOps, platform engineering, SRE, or similar infrastructure roles focused on ML systems.
  • Hands-on experience with model-serving frameworks such as vLLM, TGI, Triton, or equivalent.
  • Strong experience with container orchestration and production GPU-based workloads.
  • Experience with MLOps tooling, including model registries, experiment tracking, and automated deployment pipelines.
  • Proficiency in Python and infrastructure-as-code tools such as Terraform and Helm.
  • Strong understanding of distributed systems, performance tuning, and production reliability engineering.
  • Ability to use AI coding assistants for development and debugging workflows.
  • Preferred experience with Kubeflow, MLflow, or KubeAI.
  • Preferred knowledge of GPU scheduling, CUDA or ROCm optimization, and multi-tenant inference systems.
  • Preferred experience with GPU cost optimization, early-stage startups, greenfield infrastructure, and production systems built from scratch.
  • Fluent English is required.

Benefits

  • Fully remote role for candidates working in EMEA time zones
  • ASAP start date
  • Opportunity to build foundational ML inference infrastructure from the ground up
  • Work at the intersection of distributed systems, GPU computing, and sustainable cloud architecture
  • Influence core engineering decisions and define scalable engineering best practices

Tech Stack

HelmMLflowPythonTerraform
Pragmatike

About Pragmatike

11-50 employees

Trusted by remote-first companies worldwide. Completing tech projects for startups and scaleups.