3 months ago
Remote, IndiaAny Level
Responsibilities
- Operate and optimize Kubernetes clusters, Istio service mesh, and Linux-based systems.
- Automate infrastructure workflows using Go, Python, and Shell scripting.
- Build monitoring and observability solutions with Prometheus, Grafana, and Loki.
- Troubleshoot complex networking, storage, and system performance issues.
- Participate in on-call rotations and postmortem reviews to improve system resilience.
- Partner with AI/ML teams to prepare infrastructure for model training and data pipelines.
- Lead projects from architecture through automation and intelligent monitoring while collaborating with clients and teammates.
Requirements
- Experience with Google Cloud and Terraform.
- Strong knowledge of microservices, containers including Kubernetes and Docker, and networking.
- A focus on automation, scalability, and reliability consistent with an SRE mindset.
- Hands-on experience with PKI, service mesh, and Linux systems administration.
- The successful applicant must fulfill the requirements necessary to obtain a background check.
Benefits
- Fully remote work from home with no daily office travel requirement.
- Competitive total rewards package.
- Training allowance, professional development days, training opportunities, and certification support.
- Home-office equipment including a laptop with a choice of operating system and an annual workspace budget.
- Annual wellness budget, paid vacation and sick days, and a volunteer day.
- Morning and afternoon hiring shifts are available: 4:30 AM–1 PM IST or 12:30 PM–9 PM IST.
Categories
Site Reliability