2 months ago
Base Salary
$231k - $340k/yr
Responsibilities
- Own the data platform architecture and technical direction, including reusable frameworks and build-versus-buy decisions.
- Build and operate streaming, batch, CDC, and third-party ingestion systems with safe schema evolution and Snowflake delivery.
- Own workflow orchestration for scheduling, retries, backfills, and dependency management.
- Build transformation, compute, and self-service pipeline frameworks for product engineers, data engineers, and analysts.
- Design and operate stream-processing infrastructure for real-time product features, alerting, and reporting.
- Build data quality, observability, lineage, cataloging, discovery, and trust systems.
- Create tooling for PII classification, masking, retention, access control, and multi-region data residency.
- Set technical standards through design reviews and documentation, mentor teammates, and help build the data platform team.
Requirements
- 10+ years building and operating production data infrastructure and owning systems relied on by other teams.
- Deep experience with cloud data warehouses, especially Snowflake, including performance tuning and cost management.
- Hands-on experience building CDC and streaming pipelines with technologies such as Kafka, Debezium, Flink, or Spark Streaming.
- Experience with managed ingestion tools such as Fivetran or Airbyte and sound build-versus-buy judgment.
- Strong experience operating workflow orchestration platforms such as Temporal, Airflow, or Dagster at scale.
- Strong programming skills in Python and advanced SQL.
- Experience building frameworks or internal tools used by other engineers.
- Practical experience with data quality, observability, lineage, and schema evolution.
- Working knowledge of PII governance, masking, access control, retention, and data residency in regulated environments.
- Familiarity with Azure, AWS, GCP, Kubernetes, and infrastructure-as-code tools such as Terraform or Pulumi.
- Experience with dbt, lakehouse architectures, Iceberg, Delta Lake, Trino, multi-tenant secure platforms, or AI-product data infrastructure is preferred.
Benefits
- The role is based in San Francisco, California or New York, New York.
- The position offers the opportunity to join as one of the first hires on Harvey’s central data platform team and help shape the team and technical direction.
- Harvey is an equal opportunity employer and provides reasonable accommodations for applicants with disabilities.
Tech Stack
AirbyteAmazon RedshiftApache AirflowApache FlinkApache KafkaAWSAzureDatabricksdbtGoogle BigQueryGoogle Cloud PlatformKubernetesPythonSnowflakeSQLTerraform
Categories
BackendData Engineering
About Harvey
Harvey builds domain-specific generative AI for legal and other professional services, delivered as an enterprise platform and APIs to automate contract analysis, due diligence, compliance, and litigation workflows. Founded in 2022 and headquartered in San Francisco, it sells to law firms and corporate legal departments on enterprise agreements; investors include Sequoia Capital and OpenAI. Customers include multiple Am Law 100 firms and Fortune 500 companies’ legal teams.
