5 months ago
Base Salary
$100k - $160k/yr
Responsibilities
- Respond to automated alerts and customer-reported production issues.
- Triage, diagnose, and resolve incidents with an emphasis on permanent fixes.
- Create and maintain incident response playbooks and postmortem processes.
- Coordinate with customer success managers and key account stakeholders during incidents.
- Design and instrument telemetry, logging, alerting, dashboards, and health metrics across the serverless AWS stack.
- Identify recurring failure patterns, improve codebase resilience, and reduce operational toil through automation.
- Contribute to the full-stack codebase and assess reliability risks before new features launch.
- Make data-driven recommendations for investments in system stability.
Requirements
- At least 2 years of software engineering experience, including meaningful time in reliability, platform, or production-facing roles.
- Strong debugging skills and experience tracing distributed-system failures using logs, traces, and metrics.
- Hands-on experience with AWS, including Lambda, SQS, RDS, and CloudWatch or equivalent tools.
- Ability to read and write Go, TypeScript, or similar backend languages.
- Experience building or improving observability infrastructure such as alerting, dashboards, and telemetry.
- Strong ownership of incident resolution, postmortems, and shipped fixes.
- Experience in legaltech, fintech, healthtech, or other high-sensitivity, always-on environments is a strong plus.
Benefits
- Comprehensive health, dental, and vision insurance.
- Meals in the office and regular team events.
- Competitive equity component.
- The role is based in an office environment; the posting does not specify a remote or hybrid arrangement.
Tech Stack
Categories
Site Reliability
