
Senior Site Reliability Engineer
OutSystems27 days ago
Bengaluru, IndiaSenior
Responsibilities
- Lead and onboard services and teams to Site Reliability Engineering reliability tenets.
- Establish and maintain Service Level Objectives, Service Level Agreements, and related reliability indicators.
- Design and implement scalable, reliable, secure, and cloud-native infrastructure.
- Collaborate with software development teams to improve system resilience, observability, fault tolerance, recoverability, scalability, and performance.
- Implement monitoring, alerting, logging, and tracing solutions.
- Lead incident response, resolve production issues, and conduct root-cause analyses and post-mortems.
- Automate operational tasks and accelerate incident detection and recovery using Python and Gen AI tooling.
- Communicate reliability and performance updates to stakeholders and participate in a 24/7 on-call rotation.
- Foster continuous improvement, knowledge sharing, and effective collaboration across the SRE team.
Requirements
- BS/MS in Computer Science or equivalent.
- 6+ years of experience in Site Reliability Engineering managing infrastructure and services at scale.
- History of end-to-end project delivery.
- Experience managing Hadoop and Kubernetes infrastructure and related services, or equivalent experience.
- Advanced knowledge of Linux, networking, and containers.
- Proficiency in at least one high-level programming language such as Python or GoLang.
- Strong troubleshooting and debugging skills, including for complex distributed systems.
- Fluent English and excellent written and verbal communication skills.
- Understanding or hands-on experience with prompt engineering in software development.
- Familiarity with AI-native IDEs or AI assistants such as Cursor, GitHub Copilot, and Claude.
- Experience with SLOs, SLIs, and SLAs; Kubernetes and EKS; Infrastructure as Code tools; Python, Go, or Bash/Shell scripting; AWS services; and monitoring tools such as Grafana, ELK, and Prometheus is valued but not fully required.
- CKA, CKAD, or CKS certifications are valued.
Benefits
- Hybrid/onsite work arrangement in Bangalore.
- Professional Development Fund and Internal Mobility Program supporting vertical progression, lateral moves, and specialized AI skills.
- Opportunities to collaborate with a global team of experienced enterprise software professionals and mentors.
- Inclusive, diverse, and equal-opportunity workplace culture.
- Global collaboration across OutSystems offices and remote employees.
Categories
Site Reliability