13 hours ago
Bengaluru, IndiaSenior
Responsibilities
- Deploy and maintain highly available production environments across global datacenters.
- Perform proactive troubleshooting and performance analysis of internal services and cloud environments.
- Support rotating 24x7 on-call incident escalations and drive higher service availability and reliability.
- Monitor server usage, capacity, and performance and coordinate with users and vendors to resolve issues.
- Build self-healing features and automation to reduce operational effort and improve uptime.
- Collaborate with global R&D, SRE, and Cloud Services teams on deployment automation for new cloud services.
- Configure and troubleshoot web applications to ensure SLA compliance.
- Operate production environments supporting customer workloads and services.
Requirements
- 7+ years of experience in development, operations, site reliability, systems, or comparable cloud engineering.
- 7+ years of experience with Unix, Linux, or Windows operating systems.
- Experience with at least one scripting language: Python, PowerShell, Bash, or Shell Script.
- Proficiency with Ansible or comparable configuration management tools such as Chef or Puppet.
- Experience with Jenkins or a similar build automation tool.
- Experience with monitoring and logging services such as Elasticsearch, Wavefront, or Uptime.
- Experience configuring and troubleshooting web applications for SLA compliance.
- Preferred experience with database administration and troubleshooting, especially Microsoft SQL Server.
- Preferred experience with cloud infrastructure and services including vSphere products, AWS, BIG-IP, DYN, and Route53.
- Strong written and verbal communication skills.
Benefits
- The role works with a global engineering team across the US, UK, and India.
- The position includes rotating 24x7 on-call support for incident escalations.
- Omnissa offers community and environmental initiatives, matching donations to qualified nonprofits, and time off for volunteering.
Tech Stack
Categories
DevOpsSite Reliability
