Databricks

Staff Software Engineer, Lakeflow Pipelines DR

Databricks
Apply
14 hours ago
Mountain View, CA, USA or San Francisco, CA, USAStaff+
H1B sponsor

Base Salary

$192k - $260k/yr

Responsibilities

  • Design and implement distributed systems for cross-region replication and recovery of Lakeflow pipelines.
  • Build recovery capabilities for streaming tables, materialized views, checkpoints, source offsets, stateful operator state, table versions, transaction metadata, schedules, and dependencies.
  • Develop failover and failback workflows with conservative correctness guardrails.
  • Work on distributed consistency, idempotency, causal ordering, metadata reconciliation, and safe recovery.
  • Develop observability, failure-injection tests, high-fidelity recovery simulations, and game-day testing.
  • Define and execute a multi-year technical vision through incremental, production-quality deliverables.

Requirements

  • Strong software engineering skills in Java, Scala, C++, Go, Python, or a similar production language.
  • Passion for distributed systems, databases, storage systems, streaming systems, or reliability engineering.
  • Understanding of consistency, transactions, idempotency, replication, checkpointing, and data lineage.
  • Ability to define and work toward a multi-year technical vision with incremental, production-quality deliverables.
  • Eight or more years of experience working on related systems is preferred.
  • A PhD or advanced research experience in databases, distributed systems, or storage is optional.

Benefits

  • Databricks states that comprehensive benefits and perks are offered, with specific details varying by region.

Tech Stack

Categories

Databricks

About Databricks

10,000+ employees

Databricks builds a cloud-based data and AI platform centered on the lakehouse architecture, combining data engineering, analytics, and machine learning with Apache Spark, Delta Lake, and MLflow. It sells subscriptions and cloud services to enterprises that need to unify data pipelines and develop large-scale AI and analytics. Founded in 2013 by the creators of Apache Spark and headquartered in San Francisco, the company is privately held and serves organizations across many industries.

Contact me