Datavant

Senior Site Reliability Engineer

Datavant
Apply
16 days ago
Remote, United StatesSenior

Responsibilities

  • Own the lifecycle of Databricks and Snowflake platforms, including automation, workspace governance, job orchestration, and cost optimization.
  • Architect resilient, scalable, secure, and fault-tolerant infrastructure across cloud environments, including failover, autoscaling, chaos testing, and capacity planning.
  • Build and maintain platform-wide monitoring, alerting, and logging infrastructure using Datadog and other open tooling, and define SLOs and SLAs.
  • Automate deployments of data pipelines, ML workflows, and infrastructure components using GitHub Actions, Terraform, and related infrastructure-as-code tools.
  • Build patterns and tooling for inter-cloud and intra-cloud data movement across Snowflake, S3, Delta Lake, Kafka, and related platforms.
  • Develop event-driven data systems using EventBridge, SNS, SQS, and Lambda.
  • Partner with data engineers, ML engineers, data scientists, analysts, and application engineering teams to deliver a secure, self-service, production-grade platform.
  • Influence engineering-wide strategy for data platform architecture, ML enablement, and data products.

Requirements

  • 6+ years of experience in SRE, platform engineering, or DevOps roles supporting data-intensive or ML-powered applications.
  • Daily working experience with AI-assisted development tools such as Claude Code, Cursor, Copilot, or equivalent.
  • Hands-on Databricks experience covering workspace setup, cluster and job management, and CI/CD and orchestration integrations, plus Snowflake experience.
  • Deep understanding of AWS or similar cloud-native infrastructure, including VPCs, IAM, event-driven architectures, and serverless compute.
  • Expertise with observability tools, especially Datadog, and with platform-wide logging and monitoring solutions.
  • Strong command of CI/CD tooling, especially GitHub Actions, Terraform, and deployment automation for data systems.
  • Working knowledge of shell scripting and Python.
  • Experience building and supporting highly available and fault-tolerant systems.
  • Strong communication and cross-functional collaboration skills.
  • Preferred experience with DevSecOps practices, MLflow, feature stores, GPU workload orchestration, large-scale Databricks and Snowflake lakehouses, Iceberg v3, Glue, HIPAA or SOC 2 environments, Azure, multi-cloud or hybrid-cloud data platforms, and open-source infrastructure, SRE, or observability contributions.

Benefits

  • The role includes Datavant’s total rewards program; specific benefits are not detailed.
  • Post-offer health screenings and proof or completion of vaccinations may be required for client-related work, with case-by-case exemption review where applicable.
  • The posting states that Datavant is an equal opportunity employer and provides reasonable accommodations.

Tech Stack

Apache KafkaAWSAzureDatabricksDatadogGitHub ActionsMLflowPythonSnowflakeTerraform

Categories

DevOpsSite Reliability
Datavant

About Datavant

5,001-10,000 employees

Datavant builds a healthcare data collaboration platform and network that enables privacy-preserving exchange, linkage, and Release of Information across providers, payers, life sciences, and researchers. It sells software and data services for interoperability, de-identification, and compliant record retrieval, used to route more than 60 million health records among thousands of organizations. Privately held and headquartered in New York City, it reports working with 75% of the 100 largest U.S. health systems and 350+ real-world data partners.

Contact me