Harvey

Senior Software Engineer, Data Platform

Harvey
Apply
2 months ago

Base Salary

$193k - $290k/yr

Responsibilities

  • Own the data platform architecture and technical direction, including reusable frameworks and build-versus-buy decisions.
  • Build and operate ingestion across streaming, batch, CDC, and third-party connectors with safe schema evolution.
  • Land data in Snowflake with reliable freshness, completeness, and cost characteristics and define handoffs to Analytics Engineering.
  • Own orchestration capabilities including scheduling, retries, backfills, and dependency management.
  • Build transformation, compute, stream-processing, and self-serve pipeline frameworks for product engineers, data engineers, and analysts.
  • Develop data quality, observability, lineage, cataloging, discovery, alerting, and anomaly-detection capabilities.
  • Create tooling for PII classification, masking, retention, access control, and multi-region data residency.
  • Set technical standards through design reviews, documentation, and mentorship while helping build the data platform team.

Requirements

  • 5+ years of experience building and operating production data infrastructure that other teams depend on.
  • Deep experience with cloud data warehouses, especially Snowflake, including performance tuning and cost management.
  • Hands-on experience building CDC and streaming pipelines with technologies such as Kafka, Debezium, Flink, or Spark Streaming.
  • Experience with managed ingestion tools such as Fivetran or Airbyte and sound build-versus-buy judgment.
  • Strong experience operating workflow orchestration platforms such as Temporal, Airflow, or Dagster at scale.
  • Strong Python programming skills and advanced SQL proficiency.
  • Experience building frameworks or internal tools used by other engineers.
  • Practical experience with data quality, observability, lineage, and schema evolution for high-availability systems.
  • Working knowledge of data governance in regulated environments, including PII classification, masking, access control, retention, and data residency.
  • Familiarity with Azure, AWS, GCP, Kubernetes, and infrastructure-as-code tools such as Terraform or Pulumi.
  • Preferred: experience with dbt, lakehouse architectures, Iceberg, Delta Lake, Trino, multi-tenant secure platforms, AI-product data infrastructure, or early/founding data platform roles.

Benefits

  • The role is based in San Francisco, CA or New York, NY.
  • The position offers the opportunity to be an early hire on the central data platform team and help shape its technical direction and team.
  • Compensation is listed separately as a base salary range of $193,400–$290,000 USD.

Tech Stack

AirbyteAmazon RedshiftApache AirflowApache FlinkApache KafkaAWSAzureDatabricksdbtGoogle BigQueryGoogle Cloud PlatformKubernetesPythonSnowflakeSQLTerraform

Categories

Data Engineering
Harvey

About Harvey

1,001-5,000 employees

Harvey builds domain-specific generative AI for legal and other professional services, delivered as an enterprise platform and APIs to automate contract analysis, due diligence, compliance, and litigation workflows. Founded in 2022 and headquartered in San Francisco, it sells to law firms and corporate legal departments on enterprise agreements; investors include Sequoia Capital and OpenAI. Customers include multiple Am Law 100 firms and Fortune 500 companies’ legal teams.

Contact me