
Senior Site Reliability Engineer
OutSystems27 days ago
Bengaluru, IndiaSenior
Responsibilities
- Lead and onboard services and teams to reliability engineering practices.
- Establish and maintain Service Level Objectives and Service Level Agreements.
- Design and implement scalable, reliable, secure, and cloud-native infrastructure.
- Collaborate with software development teams to improve system resilience, observability, fault tolerance, recoverability, scalability, and performance.
- Implement monitoring, alerting, logging, and tracing solutions.
- Lead incident response, resolve production issues, and conduct root-cause analyses and post-mortems.
- Automate operational tasks with an emphasis on rapid incident detection and recovery.
- Develop mission-critical automation and tools in Python using generative AI tooling.
- Communicate system reliability and performance updates to stakeholders.
- Participate in an on-call rotation providing 24/7 production support.
- Foster continuous improvement and knowledge sharing across the SRE team.
Requirements
- BS/MS in Computer Science or equivalent.
- 6+ years of experience in Site Reliability Engineering managing infrastructure and services at scale.
- History of end-to-end project delivery.
- Experience managing Hadoop and Kubernetes infrastructure and related services, or equivalent experience.
- Advanced knowledge of Linux, networking, and containers.
- Proficiency in at least one high-level programming language such as Python or GoLang.
- Strong troubleshooting and debugging skills, particularly for complex distributed systems.
- Fluency in English and excellent written and oral communication skills.
- Understanding of or hands-on experience with prompt engineering in software development.
- Familiarity with AI-native IDEs or AI assistants such as Cursor, GitHub Copilot, and Claude.
- Experience establishing, monitoring, and improving SLOs, SLIs, and SLAs is valued.
- Experience with Kubernetes, EKS, and container orchestration is valued; CKA, CKAD, or CKS certifications are valued.
- Experience with infrastructure-as-code and automation tools such as AWS CloudFormation, Terraform, Puppet, Chef, or Spacelift is valued.
- Experience with Python, Go, Bash, shell scripting, or other automation languages is valued.
- Familiarity with AWS services such as EC2, RDS, ELB, CloudFront, and Lambda is valued.
- Experience with Grafana, the ELK stack, Prometheus, or comparable monitoring tools is valued.
- Strong understanding of resilient and fault-tolerant system design is valued.
Benefits
- Hybrid/onsite work arrangement in Bangalore.
- Professional Development Fund for specialized learning and AI skills.
- Internal Mobility Program supporting vertical progression and lateral moves.
- Structured professional development programs.
- Global collaboration with experienced technical mentors and colleagues.
- Inclusive and diverse culture with equal opportunity employment practices.
Categories
Site Reliability