
Lead Site Reliability Engineer
Air Arabia20 hours ago
Pune, IndiaStaff+
Responsibilities
- Lead Site Reliability Engineering initiatives for mission-critical Java microservices, monolithic applications, and airline Passenger Service Systems.
- Define and improve SLAs, SLOs, error budgets, reliability reporting, system availability, performance, scalability, and resilience.
- Lead production incident response, root cause analysis, post-incident reviews, on-call escalation, and preventive remediation.
- Drive architectural improvements using caching, distributed-system design, monitoring, logging, observability, and automation.
- Own CI/CD pipeline standards, GitOps practices, infrastructure automation strategy, containerization standards, and Kubernetes orchestration strategy.
- Collaborate with software engineering, infrastructure, and cross-functional teams on application design, production readiness, capacity planning, and release readiness.
- Mentor team members and provide technical guidance on incident management, production support, troubleshooting, and operational excellence.
- Evaluate reliability tools and technologies and represent SRE in planning and management reviews.
Requirements
- Bachelor’s degree in Computer Engineering, Computer Science, or Information Technology.
- 6–8 years of experience in software engineering, SRE, or Java application production support.
- Strong expertise in Java, Spring Boot, enterprise production troubleshooting, distributed systems, system architecture, scalability, high availability, and resilient application design.
- Mandatory hands-on experience with Docker and Kubernetes for containerization, orchestration, and production deployment.
- Experience designing, developing, and supporting both microservices and monolithic application architectures.
- Experience defining CI/CD standards, GitOps practices, and Git-based workflows across multiple teams.
- Strong SQL knowledge, including database design, optimization, and performance tuning; Oracle Database experience is advantageous.
- Experience with monitoring, logging, and observability tools such as Prometheus, Grafana, Elasticsearch, or Datadog.
- Hands-on experience with caching technologies such as Redis and messaging platforms.
- JBoss Application Server or similar enterprise Java application server experience is advantageous.
- Airline, aviation, or travel industry experience, particularly with Passenger Service Systems, is preferred.
- Fluent English and proficiency in Microsoft Office are required.
Tech Stack
Categories
Site Reliability