Sunset

Platform / Site Reliability Engineer

Sunset
Apply
4 days ago

Responsibilities

  • Establish a baseline for the platform, workloads, reliability, ownership, toil, recovery, costs, and technical controls.
  • Build reusable infrastructure-as-code modules, runtime templates, deployment workflows, environment contracts, and operational tooling.
  • Create supported platform paths for customer-facing services, asynchronous and batch jobs, data pipelines, and model-backed workloads.
  • Improve deploy safety, workload visibility, backup and recovery, incident response, replay, rollback, and durable remediation.
  • Define service and pipeline objectives, ownership, escalation, and recovery paths with engineering teams.
  • Build self-service for infrastructure, environments, access, deployment, debugging, and recovery without becoming a central approval queue.
  • Make cloud and vendor costs understandable by service and workload and improve efficiency within reliability and security constraints.
  • Partner with Security on cloud identity, secrets, isolation, audit logging, vulnerability response, incident readiness, and automated control evidence.
  • Support employees and contractors through bounded access, safe environments, release controls, documentation, and timely removal of authority.
  • Use AI tools deeply in platform engineering and operations while verifying generated code, plans, queries, state changes, and incident conclusions.

Requirements

  • Personally owned production cloud infrastructure and delivery or reliability systems across multiple services, including asynchronous, batch-data, or model-backed workloads.
  • Strong software engineering skills across application, platform, and infrastructure code, with experience operating the result in production.
  • Ability to reason across user impact, dependencies, state, telemetry, incidents, recovery, and durable remediation.
  • Experience building adopted paved roads that make engineering work easier.
  • Understanding of reliability models for both long-running services and high-volume or scheduled workloads.
  • Ability to balance delivery speed, least privilege, isolation, recovery, developer experience, and unit cost.
  • Effectiveness in an early-stage environment where ownership and a trustworthy baseline must be established.
  • Calm incident leadership, clear communication, and commitment to durable system improvements.
  • Fluency with modern AI engineering tools and ability to verify generated infrastructure, queries, code, and operational conclusions.
  • Bonus experience as an early platform or SRE hire, with AWS, Terraform, container runtimes, workflow orchestration, observability systems, high-volume data processing, model serving, evaluation jobs, GPU workloads, machine-learning platforms, developer environments, preview systems, CI/CD, progressive delivery, internal developer platforms, replayable pipelines, backup and restore, disaster recovery, capacity planning, cloud-cost allocation, SOC 2 controls, or enterprise customer requirements.

Tech Stack

Categories

DevOpsSite Reliability
Sunset

About Sunset

11-50 employees

We help tech companies shut down. From state withdrawals to liquidations, we save founders thousands of dollars, hundreds of hours, and countless headaches when it comes to winding down their operations.