6 days ago
Pune, IndiaSenior
Responsibilities
- Design, build, administer, and optimize highly available Linux production infrastructure for backup and disaster recovery services.
- Manage production Kubernetes clusters for reliability, scalability, security, and efficient resource utilization.
- Build modular Infrastructure as Code using Terraform or Pulumi.
- Design and maintain CI/CD pipelines with GitHub Actions, Jenkins, ArgoCD, or equivalent tools.
- Develop infrastructure automation using Shell, Python, or Go.
- Implement monitoring, logging, observability, security controls, and audit logging.
- Lead incident response, on-call support, root cause analysis, postmortems, alert tuning, and operational improvements.
- Collaborate with software engineering, SRE, and platform teams to improve deployments, reliability, and production readiness.
- Troubleshoot complex Linux, networking, storage, and infrastructure issues in production.
Requirements
- 8–12 years of experience in DevOps, Platform Engineering, Linux Administration, or Site Reliability Engineering.
- Strong Linux system administration expertise, including performance tuning, kernel parameters, storage and filesystem management, process management, networking, and troubleshooting.
- Hands-on experience managing Kubernetes clusters in production.
- Strong knowledge of Infrastructure as Code using Terraform or Pulumi.
- Experience designing and maintaining CI/CD pipelines with GitHub Actions, Jenkins, ArgoCD, or equivalent tools.
- Experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, or ELK.
- Strong scripting skills in Shell, Python, or Go.
- Experience with production incidents, on-call support, alert tuning, and operational excellence.
- Strong understanding of DNS, load balancers, firewalls, VPCs, and hybrid networking.
- Experience implementing security practices including Vault, RBAC, secrets management, vulnerability scanning, and audit logging.
- Preferred experience with OpenStack, including Nova, Swift, Neutron, and Cinder.
- Preferred experience managing AWS, Azure, GCP, and private cloud workloads.
- Preferred exposure to distributed infrastructure with 5,000+ nodes, storage platforms, backup infrastructure, disaster recovery, FinOps, storage tiering, capacity planning, chaos engineering, resilience testing, PXE, MaaS, or Ironic.
- Ability to read and troubleshoot Go-based services and contribute to automation tooling.
Benefits
- Pune location with a hybrid/onsite work arrangement.
- Equal employment opportunity is provided to all employees and applicants.
Tech Stack
AWSAzureDatadogGitHub ActionsGoGoogle Cloud PlatformGrafanaJenkinsKubernetesLinuxOpenStackPrometheusPythonSwiftTerraformVault
Categories
About Kaseya
Kaseya is the leading global provider of AI-powered IT management and cybersecurity software. Kaseya delivers a unified technology platform to manage infrastructure, secure endpoints, back up critical data, and streamline operations for more than 40,000 MSP and SMB customers around the globe. To learn more, visit www.kaseya.com.
