DeepJudge

Site Reliability Engineer

DeepJudge
Apply
18 days ago
Zürich, SwitzerlandMid Level

Responsibilities

  • Own the availability, latency, and performance of production search, retrieval, ingestion, and data-store environments.
  • Manage cloud infrastructure across public and private clouds, including networking, identity and access, compute, storage, capacity, quotas, and cost.
  • Improve deployment and infrastructure through GitOps-based configuration, infrastructure-as-code, and safe, recoverable production changes.
  • Define and instrument SLOs and error budgets and build dashboards, metrics, and alerting for actionable observability.
  • Reduce operational toil through automation and internal tooling written in Python and/or Go.
  • Provision and standardize new production environments to a defined reliability bar.
  • Investigate live incidents, harden systems against recurring failures, and feed reliability learnings back into product engineering and the platform.

Requirements

  • At least 3 years of experience in SRE, production or platform engineering, or infrastructure-heavy software engineering.
  • Production experience operating container-orchestrated systems and comfort with at least one major public cloud.
  • Solid experience with cloud infrastructure, capacity, quotas, and cost management.
  • Hands-on observability experience with metrics, logs, and traces, including root-cause analysis of live incidents.
  • Fluency with infrastructure-as-code and GitOps.
  • Ability to code in Python and/or Go for automation and tooling and to read and patch service code.
  • Sound production judgment and the ability to document work clearly for other engineers.
  • Preferred experience includes large-scale search or distributed data systems, database operations, dedicated or enterprise-deployed software, or AI and enterprise-search infrastructure in regulated domains.

Benefits

  • Ownership of the reliability of systems used by leading law firms at significant scale.
  • Hands-on work with distributed search, container orchestration across multiple clouds, infrastructure-as-code, and GitOps.
  • Opportunity to establish reliability practices, SLOs, incident processes, and automation from a strong foundation.
  • Collaborative engineering team based in Zurich.
  • Equal opportunity employer committed to an inclusive, high-performance culture.

Tech Stack

Categories

Site Reliability
DeepJudge

About DeepJudge

51-200 employees
Contact me