
Senior Site Reliability Engineer
Air Arabia20 hours ago
Pune, IndiaSenior
Responsibilities
- Troubleshoot and resolve complex production incidents and system performance issues across airline reservation and enterprise applications.
- Perform root cause analysis for application failures, outages, and recurring incidents and implement corrective and preventive actions.
- Analyze defects, implement fixes where appropriate, and provide technical recommendations to development teams.
- Monitor application health, availability, and operational metrics to meet reliability and uptime targets.
- Improve monitoring, alerting, logging, dashboards, and observability capabilities.
- Support CI/CD pipelines, deployment automation, and release management activities.
- Manage containerized application environments using Docker and Kubernetes.
- Participate in incident management, on-call support, and problem management.
- Review application logs, database performance, and system integrations to identify reliability risks.
- Develop automation, operational runbooks, standard operating procedures, and technical documentation.
- Support compliance with IT governance, cybersecurity, change management, and operational standards.
- Optimize application performance, infrastructure utilization, deployment processes, and operational workflows.
Requirements
- Bachelor’s degree in Computer Science, Software Engineering, Information Technology, or an equivalent discipline.
- 4–7 years of experience supporting Java-based enterprise applications in production environments.
- Strong hands-on expertise in Java and Spring Boot.
- Experience with both microservices and monolithic applications.
- Strong troubleshooting, debugging, analytical, and problem-solving skills.
- Basic scripting and automation knowledge; Python experience is preferred.
- Strong SQL knowledge and experience with database troubleshooting and performance analysis.
- Familiarity with Prometheus, Grafana, Elasticsearch, and Datadog.
- Experience with CI/CD tools and deployment pipelines such as Jenkins and GitOps practices.
- Hands-on enterprise production experience with Docker and Kubernetes is mandatory.
- JBoss application server experience is advantageous.
- Experience in airline, travel technology, reservation systems, or high-availability enterprise environments is preferred.
- Fluent English and proficiency in Microsoft Office.
Tech Stack
Categories
Site Reliability