Sunset

Platform / Site Reliability Engineer

Sunset
Apply
2 months ago

Responsibilities

  • Establish a baseline for the platform, workloads, reliability, ownership, toil, recovery, costs, and technical controls.
  • Build reusable infrastructure-as-code modules, runtime templates, deployment workflows, environment contracts, and operational tooling.
  • Create supported platform paths for customer-facing services, asynchronous and batch jobs, data pipelines, and model-backed workloads.
  • Improve deploy safety, workload visibility, backup and recovery, incident response, replay, rollback, and durable remediation.
  • Define service and pipeline objectives, ownership, escalation, and recovery paths with engineering teams.
  • Build self-service for infrastructure, environments, access, deployment, debugging, and recovery without becoming a central approval queue.
  • Make cloud and vendor costs understandable by service and workload and improve efficiency within reliability and security constraints.
  • Partner with Security on cloud identity, secrets, isolation, audit logging, vulnerability response, incident readiness, and automated control evidence.
  • Support employees and contractors through bounded access, safe environments, release controls, documentation, and timely removal of authority.
  • Use AI tools deeply in platform engineering and operations while verifying generated code, plans, queries, state changes, and incident conclusions.

Requirements

  • Personally owned production cloud infrastructure and delivery or reliability systems across multiple services, including asynchronous, batch-data, or model-backed workloads.
  • Strong software engineering skills across application, platform, and infrastructure code, with experience operating the result in production.
  • Ability to reason across user impact, dependencies, state, telemetry, incidents, recovery, and durable remediation.
  • Experience building adopted paved roads that make engineering work easier.
  • Understanding of reliability models for both long-running services and high-volume or scheduled workloads.
  • Ability to balance delivery speed, least privilege, isolation, recovery, developer experience, and unit cost.
  • Effectiveness in an early-stage environment where ownership and a trustworthy baseline must be established.
  • Calm incident leadership, clear communication, and commitment to durable system improvements.
  • Fluency with modern AI engineering tools and ability to verify generated infrastructure, queries, code, and operational conclusions.
  • Bonus experience as an early platform or SRE hire, with AWS, Terraform, container runtimes, workflow orchestration, observability systems, high-volume data processing, model serving, evaluation jobs, GPU workloads, machine-learning platforms, developer environments, preview systems, CI/CD, progressive delivery, internal developer platforms, replayable pipelines, backup and restore, disaster recovery, capacity planning, cloud-cost allocation, SOC 2 controls, or enterprise customer requirements.

Tech Stack

Categories

DevOpsSite Reliability
Sunset

About Sunset

51-200 employees

Sunset provides software-enabled services that help tech companies wind down operations, handling state withdrawals, liquidations, and related compliance to reduce cost and complexity. It also offers tools that help businesses unlock new revenue streams beyond shutdowns. Founded in 2023 and headquartered in New York, the privately held company sells fee-based services and tooling to founders and operators seeking an efficient, compliant exit.

Contact me