2 hours ago
Remote, WorldwideStaff+
Base Salary
$214k - $309k/yr
Responsibilities
- Establish service-quality strategies by defining service-level indicators and error budgets.
- Operate, optimize, and improve large-scale production systems through maintainable code, reliability practices, on-call participation, and incident response.
- Improve customer and internal user experience through performance engineering, observability adoption, and automation.
- Apply technical judgment to ambiguous problems, prototypes, hardening decisions, and tradeoffs.
- Raise the team’s technical bar through design feedback, code review, mentoring, and clear written communication.
- Use AI-powered development tools across planning, implementation, review, testing, and iteration while maintaining independent judgment.
Requirements
- Experience operating and optimizing large-scale distributed systems in production, including observability and self-healing techniques.
- Expertise managing production services in cloud environments such as AWS, Azure, or GCP.
- Experience with monitoring and observability tools such as Datadog, Grafana, Prometheus, and OpenTelemetry.
- Experience operating production systems in a high-growth startup environment.
- Familiarity with declarative production-infrastructure management and modern Infrastructure as Code tools such as Kubernetes and Terraform.
Benefits
- Remote-first team with a flexible-first culture.
- Equity stock options.
- 401(k) with a 5% company match and immediate vesting.
- Unlimited paid time off.
- Medical, dental, and vision insurance.
- Generous parental leave.
- Life insurance and disability benefits.
- Monthly remote-work stipend.
