Qutwo

Senior Site Reliability Engineer

Qutwo
Apply
7 hours ago
Helsinki, FinlandSenior

Responsibilities

  • Build and operate a portable, multi-cloud platform based on Kubernetes.
  • Design Kubernetes infrastructure covering cluster operations, networking, storage, troubleshooting, and workload execution.
  • Build metrics, logging, tracing, and alerting systems and define service-level objectives and incident practices.
  • Automate reproducible and reviewable environments with infrastructure as code and GitOps practices.
  • Own scheduling, autoscaling, and cost efficiency for distributed ML training and inference workloads, including Ray clusters on Kubernetes.
  • Implement secrets management, network policy, workload identity, and supply-chain security.
  • Build paved roads for CI/CD and deployment and document infrastructure decisions clearly.

Requirements

  • At least 7 years of experience building and operating production systems, including several years in an SRE, platform, or infrastructure role.
  • Deep hands-on Kubernetes experience, including cluster operations, networking, storage, and troubleshooting.
  • Experience operating systems across more than one cloud provider and understanding the associated trade-offs.
  • Strong infrastructure-as-code and automation skills, such as Terraform and Helm, with fluency in Python or Go for tooling.
  • Strong understanding of distributed systems, failure modes, and reliability-target trade-offs.
  • Experience applying security principles involving trust boundaries, identity, secrets, and software supply chains.
  • Comfort using coding agents and critically reviewing their output, including infrastructure changes.
  • Fluent English communication and the ability to work effectively in agile, cross-functional teams.
  • Preferred experience includes production Ray or distributed compute operations, GPU workload scheduling and optimization, ML pipelines, formal security or compliance requirements such as ISO 27001 or SOC 2, cloud cost management and FinOps, or familiarity with quantum computing concepts.

Categories

DevOpsSite Reliability
Qutwo

About Qutwo

51-200 employees

Qutwo builds enterprise AI and quantum-computing solutions, combining its Qutwo OS, ML pipelines for neural network compression, and security-focused SaaS services for industrial optimization, simulation, and analytics. The company partners with enterprises to develop classical, hybrid, and quantum systems and offers software plus applied research engagements. It is privately held, Europe-based, and was co-founded by pioneers behind IQM and Silo AI, with backing from founders of Skype, Hugging Face, Quantinuum, and Supercell.

Contact me