4 hours ago
Remote, IrelandStaff+
Responsibilities
- Own the reliability posture of production services, including availability, latency, capacity, efficiency, performance, monitoring, and alerting.
- Define, instrument, and operate against SLIs and SLOs, using error budgets to guide engineering priorities.
- Identify reliability risks, reduce recurring incident classes and operational toil, and drive mitigation work to completion.
- Improve production detection, response, recovery, failure-domain design, and rollback safety.
- Participate in on-call and lead incident response when production services are degraded.
- Write post-mortems, identify root causes, and oversee follow-up work.
- Write, configure, test, review, document, and deploy code that improves service reliability.
- Orchestrate complex changes across systems and services, including design changes, technical decisions, migrations, and upgrades.
- Lead debugging, troubleshooting, and analysis of service architecture and design.
- Improve engineering quality through code review, design feedback, and mentorship.
- Drive cross-team projects from conception to completion, including milestones, task tracking, timelines, stakeholder updates, and stability-impact communication.
Requirements
- 8+ years of related engineering experience, with substantial experience in reliability, infrastructure, or platform engineering.
- Demonstrated accountability for production systems, including pager and service-failure ownership.
- Strong software engineering fundamentals and experience building and shipping production code.
- Experience defining and operating against SLIs and SLOs and using error budgets.
- Depth in production operations, including incident command, post-mortem analysis, capacity planning, and observability.
- Track record of preventing incident recurrence and reducing operational toil.
- Experience driving changes across multiple teams and building alignment without formal authority.
- Experience improving engineers through code review, design feedback, and mentorship.
- Experience with large-scale distributed systems in a cloud environment.
- Preferred familiarity with infrastructure-as-code, container orchestration, GitOps-style delivery, multi-region architecture, failure-domain design, regional expansion, chaos engineering, or resilience validation.
Benefits
- Remote role based in Ireland.
- Occasional travel may be required for project or team in-person meetings.
- Competitive pay, generous time off, parental and wellness leave, healthcare, and a retirement savings program.
- Volunteering and donation support through Twilio community programs.
Categories
Site Reliability
About Twilio
Twilio builds a cloud customer-engagement platform centered on programmable communications APIs for SMS, voice, WhatsApp, email, and verification, plus Segment CDP and the Flex contact center. It sells usage-based and subscription SaaS to developers and enterprises to embed messaging, calling, authentication, and data-driven marketing into apps. Founded in 2008 and headquartered in San Francisco, Twilio is a public company listed on the NYSE.
