8 hours ago
Zürich, SwitzerlandStaff+
Responsibilities
- Build and manage cloud compute infrastructure for large-scale research workloads.
- Design and optimize distributed workflow scheduling for performance and cost efficiency.
- Develop recovery safeguards for worker interruptions, application crashes, and partial results while preventing duplicate work.
- Improve monitoring, logging, and diagnostic tools for identifying and resolving system failures.
- Partner with researchers and engineers to develop platform capabilities and shape infrastructure architecture.
Requirements
- 5 to 10+ years of relevant experience operating production infrastructure or services.
- Strong Python and Linux skills, including scripting, process management, and debugging.
- Experience with AWS or another major cloud platform covering compute, storage, networking, and access management.
- Understanding of distributed systems concepts including queues, timeouts, retries, and idempotency.
- Experience investigating production issues and balancing reliability, complexity, and cost.
- Familiarity with Terraform or similar infrastructure-as-code tools and monitoring platforms such as CloudWatch, Prometheus, or Grafana.
- Experience with virtualization, VM images and snapshots, desktop application automation, batch scheduling, experiment orchestration, or sandboxing untrusted code is a plus.
Benefits
- Competitive compensation and equity.
- Health, dental, and vision coverage.
- Visa sponsorship and relocation support.
- Full-time, on-site work in Zürich, Switzerland.
