Pythian

Site Reliability Engineer

Pythian
Apply
2 months ago
Hyderābād, IndiaSenior

Responsibilities

  • Design, deploy, operate, and optimize large-scale distributed systems across compute, storage, networking, and AI/ML environments.
  • Operate Kubernetes clusters, Istio service mesh, and Linux-based systems.
  • Automate workflows using Go, Python, and Shell scripting.
  • Build monitoring and observability solutions with Prometheus, Grafana, and Loki.
  • Troubleshoot networking, storage, and system performance issues.
  • Participate in on-call rotations and postmortem reviews to improve system resilience.
  • Partner with AI/ML teams to prepare infrastructure for model training and data pipelines.
  • Lead projects from architecture through automation and intelligent monitoring.

Requirements

  • Experience with Google Cloud and Terraform.
  • Strong knowledge of Kubernetes, Docker, networking, and containerized systems.
  • Hands-on experience with PKI, service mesh, and Linux systems administration.
  • A focus on automation, scalability, and reliability in an SRE environment.

Benefits

  • Fully remote work from home with no daily office travel requirement.
  • Morning and afternoon shift options are available.
  • Substantial training allowance, professional development days, training opportunities, and certification support.
  • Company-provided home-office equipment, including a laptop with a choice of operating system, plus an annual workspace budget.
  • Annual wellness budget, paid vacation and sick days, and a paid day off for volunteering.
  • Competitive total rewards package.

Tech Stack

AWSDockerGoGoogle CloudGrafanaIstioKubernetesLinuxPrometheusPythonSnowflakeTerraform

Categories

Site Reliability
Pythian

About Pythian

501-1,000 employees
Contact me