1 month ago
Pune, IndiaMid Level / Senior
Responsibilities
- Provision, configure, harden, and maintain Azure compute, Kubernetes, networking, database, and identity infrastructure.
- Own monitoring, alerting, dashboards, incident response, root cause analysis, post-incident reviews, and on-call coverage for a 99.9% uptime commitment.
- Implement vulnerability management, secrets management, network security controls, and certificate lifecycle management.
- Configure identity and SSO integrations and troubleshoot authentication issues.
- Manage distributed networking, connectivity, infrastructure usage, capacity planning, backups, and disaster recovery.
- Build and maintain infrastructure-as-code and CI/CD automation for provisioning and deployment.
- Support customer-facing validation, UAT, go-live sign-off, documentation, runbooks, and knowledge transfer.
Requirements
- 4–8 years of experience in Site Reliability Engineering, DevOps, or cloud infrastructure roles, including significant hands-on Azure experience.
- Hands-on experience with AKS, Virtual Machines, VNet, NSG, Load Balancer, Application Gateway, Azure Database for PostgreSQL/MySQL, Key Vault, Azure Monitor, and Log Analytics.
- Production experience administering and troubleshooting Kubernetes workloads.
- Experience with Terraform, Bicep, or ARM templates and scripting with Python, Bash, or PowerShell.
- Understanding of DNS, TLS/SSL, load balancing, firewalls, and NSGs.
- Experience with incident management, on-call rotations, SLA-driven operations, and customer stakeholder communication.
- Microsoft Certified: DevOps Engineer Expert Azure DevOps certification is required.
- Experience with OIDC/SAML identity federation, Keycloak, Okta, MQTT or other industrial data protocols, and manufacturing, industrial, or IoT customer environments is preferred.
- Additional preferred certifications include Microsoft Certified: Azure Administrator Associate, Azure Solutions Architect Expert, and Certified Kubernetes Administrator.
Benefits
- Participation in an on-call rotation for continuous coverage and incident response outside standard working hours.
- Work on production infrastructure supporting industrial AI, real-time data, and global manufacturing customers.
- High ownership, visibility, and direct impact in a collaborative, low-ego, growth-stage company.
Tech Stack
Categories
Site Reliability
