Pythian

Site Reliability Engineer

Pythian
Apply
3 months ago
Remote, IndiaAny Level

Responsibilities

  • Operate and optimize Kubernetes clusters, Istio service mesh, and Linux-based systems.
  • Automate infrastructure workflows using Go, Python, and Shell scripting.
  • Build monitoring and observability solutions with Prometheus, Grafana, and Loki.
  • Troubleshoot complex networking, storage, and system performance issues.
  • Participate in on-call rotations and postmortem reviews to improve system resilience.
  • Partner with AI/ML teams to prepare infrastructure for model training and data pipelines.
  • Lead projects from architecture through automation and intelligent monitoring while collaborating with clients and teammates.

Requirements

  • Experience with Google Cloud and Terraform.
  • Strong knowledge of microservices, containers including Kubernetes and Docker, and networking.
  • A focus on automation, scalability, and reliability consistent with an SRE mindset.
  • Hands-on experience with PKI, service mesh, and Linux systems administration.
  • The successful applicant must fulfill the requirements necessary to obtain a background check.

Benefits

  • Fully remote work from home with no daily office travel requirement.
  • Competitive total rewards package.
  • Training allowance, professional development days, training opportunities, and certification support.
  • Home-office equipment including a laptop with a choice of operating system and an annual workspace budget.
  • Annual wellness budget, paid vacation and sick days, and a volunteer day.
  • Morning and afternoon hiring shifts are available: 4:30 AM–1 PM IST or 12:30 PM–9 PM IST.

Tech Stack

AWSAzureDockerGoGoogle CloudGrafanaIstioKubernetesLinuxPrometheusPythonSnowflakeTerraform

Categories

Site Reliability
Pythian

About Pythian

501-1,000 employees
Contact me