6 months ago
Berlin, GermanySenior
Responsibilities
- Design, build, and operate cloud-native infrastructure using GKE, Kubernetes, networking, and databases.
- Own the reliability and scalability of 100+ microservices, including SLOs, autoscaling, resilience patterns, and graceful degradation.
- Lead continuous deployment practices, including rollout strategies, rollback mechanisms, and deployment observability.
- Build platform observability covering metrics, distributed tracing, structured logging, and incident diagnosis.
- Manage infrastructure as code with Terraform and Helm using peer-reviewed, version-controlled, and auditable changes.
- Advance security and compliance practices including Zero Trust architecture, Workload Identity, Vault-based dynamic secrets, network policies, ISO 27001, and SOC 2 readiness.
- Own incident response across networking, load balancing, Kubernetes, and cloud services, and drive post-incident improvements.
- Contribute to architecture and platform reviews while helping develop less-experienced engineers on the team.
Requirements
- 7+ years of total professional experience, including at least 4+ years in platform, infrastructure, or SRE roles in cloud-native environments.
- Deep Kubernetes expertise covering scheduling internals, HPA, VPA, KEDA, pod lifecycle, network policies, PodDisruptionBudgets, and multi-zone topology.
- Strong understanding of microservices operations at scale, including service mesh, resilience patterns, connection pools, graceful shutdown, and safe database migrations.
- Experience designing CI/CD pipelines for 100+ services, immutable artifact management, Workload Identity Federation, and automated rollback.
- Hands-on experience building observability platforms for metrics, logs, and distributed traces, including across asynchronous Kafka boundaries.
- Proficiency with infrastructure as code, especially Terraform and Helm.
- Programming proficiency in Golang and/or shell scripting for platform tooling.
- Proven troubleshooting experience with distributed-system failures such as latency contagion, cascading failures, connection exhaustion, and autoscaling lag.
- Collaborative working style suited to a small, high-trust engineering team.
- GitHub Actions experience is a plus.
- Experience with HashiCorp Vault and database credential rotation is a plus.
- Familiarity with GCP primitives including Workload Identity, GKE Autopilot and Standard, Cloud Armor, and VPC-native networking is a plus.
- Experience with KEDA or scheduled scaling strategies is a plus.
- Cloud cost-optimization experience and prior work in regulated FinTech or financial services are additional advantages.
Benefits
- Top-of-market compensation package including equity.
- In-person office culture across Moss’s European offices, with weekly breakfasts and Friday demos.
- 20 days of work-from-abroad allowance.
- 600 EUR/GBP Learning & Development budget.
- Additional local benefits, with tailored packages for interns and working students.
