4 months ago
Bengaluru, IndiaStaff+
Responsibilities
- Architect and build scalable, secure, highly available cloud-native platforms on Microsoft Azure.
- Design and manage infrastructure automation using Terraform, GitOps, and Infrastructure as Code.
- Build CI/CD systems, deployment orchestration, environment provisioning, and self-service developer platforms.
- Design secure networking and authentication systems using Cloudflare, API gateways, private networking, and Zero Trust principles.
- Drive platform reliability, scalability, observability, security, operational excellence, and resiliency across production environments.
- Lead incident management, root cause analysis, operational governance, deployment reliability, and production readiness initiatives.
- Build AI/LLM and agentic operational workflows for troubleshooting, automation, incident response, and platform intelligence.
- Improve developer experience, platform adoption, operational efficiency, and cloud cost optimization.
- Partner with engineering, architecture, security, and product teams on platform standards and long-term technical strategy.
- Lead architecture discussions, mentor engineers, and promote automation, scalability, reliability, and platform standardization.
Requirements
- 10+ years of experience in Platform Engineering, Infrastructure Engineering, DevOps, SRE, or Cloud Engineering.
- Extensive experience managing large-scale production environments on Microsoft Azure.
- Experience with Kubernetes, container platforms, GitOps, Terraform, and cloud-native systems.
- Experience building enterprise-grade CI/CD and deployment automation platforms.
- Experience driving operational excellence, observability, reliability engineering, and infrastructure automation initiatives.
- Experience leading cross-functional technical initiatives and mentoring engineers.
- Strong hands-on coding experience in Python, Go, or other modern programming languages.
- Deep understanding of cloud infrastructure, distributed systems, networking, platform engineering, and operational models.
- Strong troubleshooting and debugging skills across infrastructure, cloud, deployment, networking, and application layers.
- Strong communication, collaboration, stakeholder management, system design, and technical leadership skills.
- Experience with AI/LLM workflows, coding agents, and intelligent automation systems is preferred.
Benefits
- Opportunity to build foundational AI-native enterprise infrastructure and systems with substantial ownership and real-world impact.
- Headquartered in Los Altos, California; no specific remote or hybrid arrangement is stated.
