
Platform Engineer (Cloud Infrastructure, AI Platform) (m/f/d)
Allianz SE2 months ago
Munich, GermanySenior
Responsibilities
- Implement and operate AKS Kubernetes infrastructure, including cluster lifecycle, networking, resource management, autoscaling, and multi-tenancy.
- Build and maintain GitHub Actions and ArgoCD CI/CD and GitOps pipelines for testing, container builds, and deployments.
- Develop Terraform and Bicep Infrastructure as Code to provision and manage Azure resources.
- Operate Azure Container Registry, artifact management, and image security scanning workflows.
- Implement observability with Azure Monitor, Application Insights, Prometheus, and Grafana, including dashboards, alerting, and distributed tracing.
- Manage Celery workers, Redis queues, and workflow orchestration patterns supporting AI agent execution.
- Implement platform security controls including network policies, pod security standards, Key Vault integration, RBAC, and private endpoints.
- Support PostgreSQL infrastructure through backup and recovery, connection pooling, and performance tuning.
- Create self-service tooling and templates for development teams.
- Diagnose infrastructure issues, perform root-cause analysis, and implement preventative improvements.
- Collaborate with Platform Architects, Backend Engineers, and ML Engineers to translate architecture into reliable infrastructure.
Requirements
- At least 5 years of professional experience in platform engineering, SRE, or DevOps roles.
- Strong Kubernetes experience covering cluster operations, networking, storage, autoscaling, and troubleshooting.
- Infrastructure as Code experience with Terraform, Bicep, or equivalent tools.
- Production experience with Azure services including AKS, ACR, Key Vault, Azure Monitor, Virtual Networks, Private Endpoints, and Azure Policy.
- Strong CI/CD experience with GitHub Actions, ArgoCD, or similar GitOps tooling.
- Proficiency in Python for automation, scripting, and tooling.
- Experience with container security, image scanning, runtime security, network policies, and least-privilege patterns.
- Experience with Prometheus, Grafana, centralized logging, and alerting configuration.
- Familiarity with Celery, Redis, or equivalent message queue patterns.
- Strong Linux systems administration and networking fundamentals.
- Strong troubleshooting, reliability, automation, communication, collaboration, and ownership skills.
- Experience supporting AI/ML infrastructure, GPU scheduling, model serving platforms, or ML pipeline orchestration is a plus.
- Service mesh experience with Istio or Linkerd is a plus.
- Experience with Databricks or similar data platform infrastructure is a plus.
- Familiarity with Temporal, Airflow, or other workflow orchestration tools is a plus.
- FinOps, regulated-environment, and cost-optimization experience is a plus.
- CKA, CKAD, Azure Administrator, or Terraform Associate certifications are a plus.
Benefits
- Professional development courses and targeted development programs.
- Global environment with international mobility and career progression opportunities.
- Work Well programs supporting employee health, wellbeing, and work-life balance.
- Full-time, permanent employment.
Tech Stack
Apache AirflowAzureDatabricksGitHub ActionsGrafanaIstioKubernetesLinuxPostgreSQLPrometheusPythonRedisTerraform