
Staff Site Reliability Engineer
Hippocratic AI10 hours ago
Responsibilities
- Design and build the GPU management and scheduling platform for approximately 30 models running across heterogeneous hardware.
- Build metrics pipelines and decision logic for GPU utilization, admission control, load shedding, and capacity management.
- Implement autoscaling that adjusts model replicas based on real-time demand and utilization.
- Develop cloud orchestration systems and operators in Python and Go.
- Architect and operate scalable, fault-tolerant, secure production systems on AWS, GCP, or Azure.
- Build infrastructure automation and deployment pipelines using Terraform and CI/CD tooling.
- Maintain monitoring, logging, and alerting for platform reliability and performance.
- Develop and enforce security and compliance policies for a healthcare AI platform.
- Diagnose infrastructure, deployment, and operational issues with engineers and research scientists.
- Mentor engineers and help raise the team's technical bar.
Requirements
- 10+ years of professional experience across site reliability, DevOps engineering, and software engineering.
- A required Computer Science degree from a top CS program.
- Strong software engineering fundamentals, including building orchestration and scheduling systems in Python and/or Go.
- Experience designing operational-metrics-driven control loops such as autoscaling, load shedding, or admission control.
- Deep experience with infrastructure automation and CI/CD, including Terraform, GitLab CI/CD, or similar tools.
- Hands-on production experience with at least one major cloud platform: AWS, GCP, or Azure.
- Strong knowledge of Docker and Kubernetes.
- Experience with monitoring and logging stacks such as ELK, Grafana, or Datadog.
- Familiarity with secrets management and security tooling such as HashiCorp Vault, AWS KMS, or Azure Key Vault.
- Strong problem-solving, independent-working, collaboration, communication, and interpersonal skills.
- Preferred experience managing GPU fleets or scheduling workloads across heterogeneous accelerators.
- Preferred familiarity with ML inference serving and model deployment tools such as Triton, KServe, or Ray Serve.
- Preferred experience with Kubernetes autoscaling internals, including HPA, VPA, custom metrics, and custom controllers.
- Preferred experience implementing HIPAA and SOC 2 compliance.
- Preferred experience operating in an HPC environment.
- A bachelor's or master's degree in Computer Science, Computer Engineering, or a related field is listed as a nice-to-have.
Benefits
- Opportunity to work on a healthcare-only, safety-focused AI platform intended to improve patient outcomes.
- Work alongside physicians, hospital leaders, AI pioneers, researchers, and experienced technology professionals.
- Join a well-funded company backed by leading healthcare and AI investors.
- Hippocratic AI is an equal opportunity employer and provides accommodations during the hiring process upon request.
Tech Stack
Categories
Site Reliability
About Hippocratic AI
Hippocratic AI builds a safety-focused large language model and AI agents for healthcare workflows, used by health systems for patient outreach, post-discharge follow-up, and chronic-care management. It licenses its platform and tools to providers to automate and scale clinical support tasks while meeting health-system requirements. Founded in 2023 and headquartered in Palo Alto, the company is privately held.