1 month ago
Sydney, AustraliaSenior
Responsibilities
- Own reliability and core platform decisions as the company scales to hundreds of millions of users.
- Improve uptime and reduce recovery time objectives across critical services.
- Orchestrate and harden GPU clusters serving millions of AI generations per day.
- Implement platform-wide metrics, tracing, alerting, and service-level objectives.
- Optimize AWS infrastructure and reduce cloud spend without sacrificing performance.
- Write code, fix root causes, and handle production incidents.
Requirements
- 5+ years of experience operating production systems at scale.
- Strong AWS experience, including infrastructure-as-code and high-scale compute.
- Experience with Kubernetes, ECS, or similar platforms.
- Deep observability and incident response experience.
- Expertise with CI/CD and deployment pipelines.
- Familiarity with TypeScript, Next.js, React, TailwindCSS, tRPC, Postgres, and Temporal.
Benefits
- Top-of-market compensation with meaningful equity upside.
- In-person work in Sydney, Australia, with global hiring.
- Visa sponsorship and relocation assistance for people moving to Sydney.
- Company card for food, coffee, tools, and other work-related needs.
- Daily team lunch and dinner at the office.
- Unlimited workspace budget.
- Compensation reviews every six months, with bonuses and raises tied to impact.
- Paid three-day work trial in Sydney, with travel and accommodation covered if needed.
Tech Stack
Categories
Site Reliability
