3 days ago
Remote, IndiaSenior
Responsibilities
- Architect, scale, and improve the reliability, availability, and performance of multi-region microservices, APIs, and authentication infrastructure across AWS and GCP.
- Design and maintain disaster recovery, multi-region failover automation, and business continuity strategies against defined RTO and RPO objectives.
- Define and enforce SLIs, SLOs, error budgets, observability, Golden Signals monitoring, and reliability practices across engineering teams.
- Lead on-call escalation, major incident management, post-incident reviews, root-cause remediation, and adherence to 99.99% availability SLAs.
- Manage production Kubernetes EKS clusters, GitOps workflows, deployment patterns, and Terraform infrastructure across multi-account, multi-region environments.
- Build FinOps and cost-optimization dashboards covering cloud spend, unit economics, resource utilization, right-sizing, tagging, and workload optimization.
- Develop production tooling, platform automation, integrations, runbooks, architecture decision records, and technical documentation in Python or Go.
- Mentor junior and mid-level engineers and lead technical discussions.
Requirements
- 8+ years of professional software engineering experience in SRE, DevOps, or platform engineering for 24/7 mission-critical distributed systems.
- Bachelor’s degree in Computer Science, Software Engineering, or an equivalent technical discipline.
- Advanced Python or Go skills for building internal SRE platforms, tools, and API integrations.
- Deep hands-on Kubernetes expertise, including EKS or GKE cluster lifecycles, networking, ingress/egress, RBAC, and GitOps tooling.
- Advanced Terraform and AWS/GCP experience across complex multi-account cloud environments, including IAM, VPC, Transit Gateway, ALB/NLB, and Route53.
- Demonstrated leadership in FinOps, cloud cost efficiency, resource optimization, cost allocation, and financial accountability for engineering teams.
- Experience designing and testing multi-region disaster recovery architectures, automated failover, and recovery-health monitoring.
- Experience defining SLI/SLOs, managing PagerDuty schedules, and optimizing production observability platforms.
- Experience operating enterprise service meshes and production ingress or proxy systems such as Istio, Linkerd, HAProxy, or NGINX.
- Ability to lead technical discussions, write architecture documents or RFCs, mentor peers, and collaborate effectively.
- Preferred experience with chaos engineering, secrets management, DevSecOps, automated vulnerability remediation, identity services, IAM, enterprise directories, or security-focused SaaS.
Benefits
- Remote-first work arrangement within India.
- Participation in assigned on-call shifts is required.
- Internal business and interview processes are conducted primarily in English, requiring fluent spoken and written English.
- Equal opportunity employment and a collaborative environment with opportunities to share ideas and grow expertise.
Tech Stack
Categories
DevOpsSite Reliability
About JumpCloud
JumpCloud builds a unified open directory and IT management platform for organizations to manage user identities, devices, and access across Windows, macOS, Linux, and cloud resources. Its SaaS suite covers single sign-on, multi-factor authentication, Cloud RADIUS, LDAP/Active Directory integration, and server/device management for IT teams. Founded in 2012 and headquartered in Louisville, Colorado, the privately held company serves customers worldwide.
