1 day ago
Zürich, SwitzerlandSenior
Responsibilities
- Build and operate cloud compute infrastructure for large-scale research workloads.
- Design and optimize distributed workflow scheduling for performance and cost efficiency.
- Develop recovery safeguards for worker interruptions, application crashes, partial results, and duplicate-work prevention.
- Improve monitoring, logging, and diagnostic tools for identifying and resolving system failures.
- Partner with researchers and engineers to develop platform capabilities and shape infrastructure architecture.
Requirements
- Typically 5 to 10 or more years of relevant infrastructure or software engineering experience.
- Strong Python and Linux skills, including scripting, process management, and debugging.
- Experience operating services on a major cloud platform across compute, storage, networking, and access management.
- Understanding of distributed systems concepts including queues, timeouts, retries, and idempotency.
- Experience investigating production issues and balancing reliability, complexity, and cost.
- Self-directed, collaborative approach with sound judgment and strong follow-through.
- Beneficial experience includes KVM/QEMU, VM images or snapshots, desktop application automation, Terraform, monitoring tools, batch scheduling, experiment orchestration, or sandboxing untrusted code.
Benefits
- Equity is offered.
- Visa sponsorship is available.
- The role is on-site in Zürich, Switzerland.
