
Lead AI Application Engineer (Infrastructure & LLMOps)
TechBiz Global GmbH3 months ago
Remote, WorldwideStaff+
Responsibilities
- Architect and operate a highly available, low-latency, cost-efficient multi-tenant AI platform across cloud and on-premises environments.
- Implement LLMOps/MLOps practices and automated deployment pipelines for models.
- Develop inference, embeddings, and RAG as-a-service capabilities with unified APIs and abstraction layers.
- Deploy and scale vector databases, feature stores, model hosting environments, Kubernetes infrastructure, and GPU orchestration.
- Build self-service portals or CLIs, templates, and blueprints for provisioning AI environments, models, and data stores.
- Optimize data retrieval for real-time AI applications and agentic workflows.
- Conduct internal workshops and create documentation to enable product squads to use the platform.
Requirements
- 8+ years of experience in platform engineering, DevOps, or Site Reliability Engineering.
- At least 2 years of experience building AI/ML infrastructure or platforms.
- Deep experience with Kubernetes, Docker, and Terraform or Pulumi.
- Experience managing workloads across AWS, Azure, GCP, and on-premises environments, including NVIDIA AI Enterprise and OpenShift.
- Hands-on experience with vLLM, TGI, or NVIDIA Triton for model serving.
- Expertise with vector databases and traditional SQL/NoSQL databases.
- High proficiency in Python and Go or Rust for platform tooling.
- Experience building Internal Developer Platforms is a major plus.