2 months ago
Responsibilities
- Architect a unified Iceberg metastore platform serving Spark, Trino, Flink, and PyIceberg.
- Define API contracts, authorization models, per-table credential vending, and integration patterns for the data platform.
- Lead the migration strategy from Hive Metastore-backed workloads, including sequencing, backward compatibility, rollback, and cross-team coordination.
- Shape the object storage abstraction layer, including bucket provisioning, access-control policies, and secure developer-facing client libraries.
- Translate regulatory requirements into preventative controls such as audit logging, access review infrastructure, data segregation, and lifecycle enforcement.
- Improve storage efficiency, snapshot retention, and data lifecycle management through automated self-service tooling.
- Lead design reviews, establish reliability, security, and developer-experience standards, and mentor senior engineers.
Requirements
- 10+ years of professional software engineering experience.
- Demonstrated experience designing, building, and operating large-scale distributed storage or data infrastructure systems.
- Deep experience with object storage such as S3 or Azure Blob, including IAM, access control policy design, lifecycle management, and petabyte-scale operations.
- Experience leading complex, multi-quarter infrastructure projects end to end, including cross-team dependency management and migrations across many consuming teams.
- Strong background in authorization and access control design for distributed data systems.
- Preferred: deep expertise in Apache Iceberg, including table format internals, the REST Catalog specification, snapshot lifecycle management, compaction, and compute engine integration.
- Preferred: experience with compliance-sensitive infrastructure and frameworks such as SOX or ICFR.
- Preferred: experience executing large-scale data migrations with sequencing, rollback, blast-radius reduction, and data-integrity validation.
- Preferred: ability to build ergonomic, well-documented abstractions that improve developer experience and reduce engineering toil.
Tech Stack
Apache FlinkApache Spark
Categories
BackendData Engineering
About Stripe
Stripe builds programmable financial services. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Headquartered in San Francisco and Dublin, the company aims to increase the GDP of the internet.