OutSystems

Senior Site Reliability Engineer

OutSystems
Apply
27 days ago
Bengaluru, IndiaSenior

Responsibilities

  • Lead and onboard services and teams to Site Reliability Engineering reliability tenets.
  • Establish and maintain Service Level Objectives, Service Level Agreements, and related reliability indicators.
  • Design and implement scalable, reliable, secure, and cloud-native infrastructure.
  • Collaborate with software development teams to improve system resilience, observability, fault tolerance, recoverability, scalability, and performance.
  • Implement monitoring, alerting, logging, and tracing solutions.
  • Lead incident response, resolve production issues, and conduct root-cause analyses and post-mortems.
  • Automate operational tasks and accelerate incident detection and recovery using Python and Gen AI tooling.
  • Communicate reliability and performance updates to stakeholders and participate in a 24/7 on-call rotation.
  • Foster continuous improvement, knowledge sharing, and effective collaboration across the SRE team.

Requirements

  • BS/MS in Computer Science or equivalent.
  • 6+ years of experience in Site Reliability Engineering managing infrastructure and services at scale.
  • History of end-to-end project delivery.
  • Experience managing Hadoop and Kubernetes infrastructure and related services, or equivalent experience.
  • Advanced knowledge of Linux, networking, and containers.
  • Proficiency in at least one high-level programming language such as Python or GoLang.
  • Strong troubleshooting and debugging skills, including for complex distributed systems.
  • Fluent English and excellent written and verbal communication skills.
  • Understanding or hands-on experience with prompt engineering in software development.
  • Familiarity with AI-native IDEs or AI assistants such as Cursor, GitHub Copilot, and Claude.
  • Experience with SLOs, SLIs, and SLAs; Kubernetes and EKS; Infrastructure as Code tools; Python, Go, or Bash/Shell scripting; AWS services; and monitoring tools such as Grafana, ELK, and Prometheus is valued but not fully required.
  • CKA, CKAD, or CKS certifications are valued.

Benefits

  • Hybrid/onsite work arrangement in Bangalore.
  • Professional Development Fund and Internal Mobility Program supporting vertical progression, lateral moves, and specialized AI skills.
  • Opportunities to collaborate with a global team of experienced enterprise software professionals and mentors.
  • Inclusive, diverse, and equal-opportunity workplace culture.
  • Global collaboration across OutSystems offices and remote employees.

Tech Stack

Apache HadoopAWSBashChefGoGrafanaKubernetesLinuxPrometheusPuppetPythonTerraform

Categories

Site Reliability
OutSystems

About OutSystems

1,001-5,000 employees
Contact me