
AI Infrastructure Engineer (GPU) - Remote EMEA
Pragmatike5 months ago
Prague, Czechia +8 moreMid Level
Responsibilities
- Build and operate production-grade model-serving infrastructure using vLLM, TGI, Triton, or equivalent frameworks.
- Design and implement blue/green and canary deployment pipelines for ML models.
- Develop autoscaling systems, multi-model serving architectures, and intelligent request-routing layers.
- Optimize GPU utilization, memory efficiency, network throughput, and model-artifact storage performance.
- Design observability systems for inference latency, throughput, GPU usage, cost metrics, and system health.
- Manage model registries and automated, reproducible model deployment pipelines.
- Own the ML systems lifecycle from development through production, including operational support and on-call responsibilities.
- Define engineering best practices and contribute to platform scalability.
Requirements
- At least 4 years of experience in MLOps, platform engineering, SRE, or similar infrastructure roles focused on ML systems.
- Hands-on experience with model-serving frameworks such as vLLM, TGI, Triton, or equivalent.
- Strong experience with container orchestration and operating GPU-based workloads in production.
- Experience with MLOps tooling including model registries, experiment tracking, and automated deployment pipelines.
- Proficiency in Python and infrastructure-as-code tools such as Terraform and Helm.
- Strong understanding of distributed systems, performance tuning, and production reliability engineering.
- Ability to use AI coding assistants to accelerate development and debugging.
- Experience with ML platforms such as Kubeflow, MLflow, or KubeAI is preferred.
- Knowledge of GPU scheduling, CUDA or ROCm optimization, or multi-tenant inference systems is preferred.
- Experience with GPU cost optimization, early-stage startups, greenfield infrastructure, and building production systems from scratch is preferred.
- Fluent English and the ability to work independently in a remote-first environment are required.
Benefits
- Fully remote role within EMEA time zones.
- Start date is ASAP.
- Opportunity to own critical infrastructure for a rapidly scaling AI-native cloud platform.
- Build foundational ML inference systems from the ground up in a well-funded startup.
- Work across distributed systems, GPU computing, and sustainable cloud architecture.
- Influence core engineering decisions and define scalable engineering best practices.
Categories
About Pragmatike
Pragmatike is a Paris-based IT services and recruiting firm connecting remote-first companies with software engineers and tech specialists worldwide. Founded in 2022, it places contractors or full-time hires and staffs teams to complete projects for startups and scaleups. The private partnership offers access to a large network of 50,000+ specialists across 60+ countries.