about 2 hours ago
Responsibilities
- Own reliability for core production systems, including deployments and incident response.
- Drive capacity planning across cloud and physical infrastructure.
- Define and improve reliability goals and operational outcomes.
- Lead observability improvements across metrics, logs, and alert quality.
- Build automation and tooling to enhance reliability and engineering velocity.
- Partner with stakeholders to align reliability priorities with roadmap goals.
- Shape database and data-processing reliability.
- Apply engineering judgment to balance delivery speed with resilience.
Requirements
- Bachelor’s degree in Computer Science, Computer Engineering, or equivalent experience.
- 10+ years of professional software engineering experience.
- Proven track record of improving reliability in distributed systems.
- Deep expertise in observability and production support practices.
- Strong networking fundamentals for distributed systems.
- Experience with distributed systems technologies like Kubernetes and Kafka.
- Experience with cloud infrastructure and operations at scale.
- Strong software engineering fundamentals in object-oriented design.
- Strong analytical, communication, and decision-making skills.
- Experience leading high-severity incident response is a plus.
Benefits
- Collaborative and welcoming work environment.
- Immediate responsibility for new technologies.
- Empowerment to own projects and make impactful decisions.
- Long-term employment based on a permanent contract.
- Attractive benefits package including medical care and life insurance.
- Annual budget for professional development ($2,000).
