3 hours ago
Base Salary
$193k - $290k/yr
Responsibilities
- Own the data platform architecture and technical direction, including build-versus-buy decisions and reusable software frameworks.
- Build and operate streaming, batch, CDC, and third-party ingestion systems with safe schema evolution and Snowflake delivery.
- Own workflow orchestration, including scheduling, retries, backfills, and dependency management.
- Build transformation, compute, and self-service pipeline frameworks for product engineers, data engineers, and analysts.
- Design and operate real-time stream-processing infrastructure for product features, alerting, and near-live reporting.
- Build data quality, observability, anomaly detection, lineage, cataloging, and discovery capabilities.
- Develop tooling for PII classification, masking, retention, access control, and multi-region data residency.
- Set technical standards through design reviews and documentation, and mentor the growing team.
Requirements
- 5+ years of experience building and operating production data infrastructure used by other teams.
- Deep experience with cloud data warehouses, with Snowflake strongly preferred and BigQuery, Databricks, or Redshift considered transferable.
- Hands-on experience building CDC and streaming pipelines with Kafka, Debezium, Flink, or Spark Streaming.
- Experience with managed ingestion tools such as Fivetran or Airbyte and with build-versus-buy connector decisions.
- Strong experience operating workflow orchestration systems such as Temporal, Airflow, or Dagster at scale.
- Strong Python programming skills and advanced SQL proficiency.
- Experience building frameworks or internal tooling used by other engineers.
- Experience with data quality, observability, lineage, and schema evolution for highly available systems.
- Working knowledge of data governance in regulated environments, including PII classification, masking, access control, retention, and data residency.
- Familiarity with Azure, AWS, GCP, Kubernetes, and infrastructure-as-code tools such as Terraform or Pulumi.
- Preferred experience with dbt, analytics engineering teams, lakehouse architectures, Iceberg, Delta Lake, Trino, multi-tenant secure platforms, AI product data infrastructure, or serving as an early data platform hire.
Benefits
- The role is based in San Francisco, California, or New York, New York.
- Harvey provides equal employment opportunity and reasonable accommodations for applicants with disabilities.
Tech Stack
AirbyteAmazon RedshiftApache AirflowApache FlinkApache KafkaApache SparkAWSAzureDatabricksdbtGoogle BigQueryGoogle Cloud PlatformKubernetesPythonSnowflakeSQLTerraform
Categories
BackendData Engineering
About Harvey
Harvey is domain-specific AI for legal and professional services. Adopted by Fortune 500 companies like AT&T, Verizon, Cox, Koch, KKR, Bridgewater, more than 100,000 lawyers across 2,400+ customers in 70 countries and over 75% of AmLaw 100 law firms rely on Harvey to advance legal expertise faster across contract analysis, due diligence, compliance, and litigation. Backed by Sequoia Capital, OpenAI, GV, Kleiner Perkins, Coatue and EQT, Harvey is the trusted partner in modernizing the legal industry.
