Hippocratic AI

Staff Site Reliability Engineer

Hippocratic AI
Apply
10 hours ago
Menlo Park, CA, USAStaff+
H1B sponsor

Responsibilities

  • Design and build the GPU management and scheduling platform for approximately 30 models running across heterogeneous hardware.
  • Build metrics pipelines and decision logic for GPU utilization, admission control, load shedding, and capacity management.
  • Implement autoscaling that adjusts model replicas based on real-time demand and utilization.
  • Develop cloud orchestration systems and operators in Python and Go.
  • Architect and operate scalable, fault-tolerant, secure production systems on AWS, GCP, or Azure.
  • Build infrastructure automation and deployment pipelines using Terraform and CI/CD tooling.
  • Maintain monitoring, logging, and alerting for platform reliability and performance.
  • Develop and enforce security and compliance policies for a healthcare AI platform.
  • Diagnose infrastructure, deployment, and operational issues with engineers and research scientists.
  • Mentor engineers and help raise the team's technical bar.

Requirements

  • 10+ years of professional experience across site reliability, DevOps engineering, and software engineering.
  • A required Computer Science degree from a top CS program.
  • Strong software engineering fundamentals, including building orchestration and scheduling systems in Python and/or Go.
  • Experience designing operational-metrics-driven control loops such as autoscaling, load shedding, or admission control.
  • Deep experience with infrastructure automation and CI/CD, including Terraform, GitLab CI/CD, or similar tools.
  • Hands-on production experience with at least one major cloud platform: AWS, GCP, or Azure.
  • Strong knowledge of Docker and Kubernetes.
  • Experience with monitoring and logging stacks such as ELK, Grafana, or Datadog.
  • Familiarity with secrets management and security tooling such as HashiCorp Vault, AWS KMS, or Azure Key Vault.
  • Strong problem-solving, independent-working, collaboration, communication, and interpersonal skills.
  • Preferred experience managing GPU fleets or scheduling workloads across heterogeneous accelerators.
  • Preferred familiarity with ML inference serving and model deployment tools such as Triton, KServe, or Ray Serve.
  • Preferred experience with Kubernetes autoscaling internals, including HPA, VPA, custom metrics, and custom controllers.
  • Preferred experience implementing HIPAA and SOC 2 compliance.
  • Preferred experience operating in an HPC environment.
  • A bachelor's or master's degree in Computer Science, Computer Engineering, or a related field is listed as a nice-to-have.

Benefits

  • Opportunity to work on a healthcare-only, safety-focused AI platform intended to improve patient outcomes.
  • Work alongside physicians, hospital leaders, AI pioneers, researchers, and experienced technology professionals.
  • Join a well-funded company backed by leading healthcare and AI investors.
  • Hippocratic AI is an equal opportunity employer and provides accommodations during the hiring process upon request.

Categories

Site Reliability
Hippocratic AI

About Hippocratic AI

201-500 employees

Hippocratic AI builds a safety-focused large language model and AI agents for healthcare workflows, used by health systems for patient outreach, post-discharge follow-up, and chronic-care management. It licenses its platform and tools to providers to automate and scale clinical support tasks while meeting health-system requirements. Founded in 2023 and headquartered in Palo Alto, the company is privately held.

Contact me