2 months ago
Hyderābād, IndiaSenior
Responsibilities
- Design, deploy, operate, and optimize large-scale distributed systems across compute, storage, networking, and AI/ML environments.
- Operate Kubernetes clusters, Istio service mesh, and Linux-based systems.
- Automate workflows using Go, Python, and Shell scripting.
- Build monitoring and observability solutions with Prometheus, Grafana, and Loki.
- Troubleshoot networking, storage, and system performance issues.
- Participate in on-call rotations and postmortem reviews to improve system resilience.
- Partner with AI/ML teams to prepare infrastructure for model training and data pipelines.
- Lead projects from architecture through automation and intelligent monitoring.
Requirements
- Experience with Google Cloud and Terraform.
- Strong knowledge of Kubernetes, Docker, networking, and containerized systems.
- Hands-on experience with PKI, service mesh, and Linux systems administration.
- A focus on automation, scalability, and reliability in an SRE environment.
Benefits
- Fully remote work from home with no daily office travel requirement.
- Morning and afternoon shift options are available.
- Substantial training allowance, professional development days, training opportunities, and certification support.
- Company-provided home-office equipment, including a laptop with a choice of operating system, plus an annual workspace budget.
- Annual wellness budget, paid vacation and sick days, and a paid day off for volunteering.
- Competitive total rewards package.
Categories
Site Reliability