about 4 hours ago
Responsibilities
- Own the design and delivery of reliability projects and features.
- Partner with critical T0/T1 services to improve scalability and reduce operational toil.
- Build and enhance systems for secure management of service configurations and secrets.
- Improve canary-based release systems and expand deployment capabilities.
- Drive reliability best practices and strengthen reliability culture across teams.
Requirements
- 10+ years of software engineering experience in service-oriented architectures.
- Experience with Ruby, Go, Terraform, and cloud platforms (AWS, GCP, or Azure).
- Ability to design and operate reliable, high-throughput, low-latency distributed systems.
- Proven experience with observability and monitoring tools like Kibana and Datadog.
- Strong communication skills for architecture decisions with cross-functional teams.
- Willingness to participate in on-call rotations and respond to issues outside normal hours.