
Sr. Site Reliability Engineer
Tiger Analytics Inc.4 months ago
Washington, DC, USASenior
Responsibilities
- Define, monitor, and maintain SLOs and SLIs for critical AI/ML services.
- Manage error budgets and balance ML feature-release velocity with production stability.
- Architect and manage Kubernetes autoscaling strategies for model training and high-volume inference.
- Ensure high availability of Vertex AI endpoints and custom inference services.
- Optimize GPU and TPU utilization for cost-efficient LLM performance.
- Support and stabilize Vertex AI Pipelines and Kubeflow ML pipelines.
- Provision and manage cloud environments with Terraform or Pulumi.
- Design deployment pipelines for application code and ML models using GitHub Actions, Cloud Build, or ArgoCD.
- Develop Python or Go automation, self-healing mechanisms, and resource-cleanup tools.
- Build monitoring and observability dashboards with Prometheus, Grafana, or Google Cloud Operations Suite.
- Participate in on-call rotations, lead technical resolution of production outages, and conduct blameless post-mortems and root-cause analysis.
Requirements
- Expert-level knowledge of Kubernetes and Docker.
- Familiarity with Kubeflow, Vertex AI, MLflow, or DVC.
- Strong proficiency in Python and Bash; knowledge of Go is a plus.
- Experience managing the reliability of data-heavy services such as BigQuery, Pub/Sub, or vector databases including Pinecone and Milvus.
- Solid understanding of VPCs, load balancers, DNS, and secure service meshes such as Istio and Anthos.
Benefits
- Significant career development opportunities in a small, fast-growing, challenging, and entrepreneurial environment.
- High degree of individual responsibility.
- On-site position in Washington, United States.
Tech Stack
Categories
DevOpsSite Reliability