4 months ago
Milpitas, CA, USAStaff+
Base Salary
$274k - $304k/yr
Responsibilities
- Define architectural standards for governed, tenant-isolated, PII-free, traceable AI training data.
- Own the Google Cloud Storage data lake layout, tiering, encryption, lifecycle rules, retention, access control, and Iceberg maintenance.
- Build and operate batch Spark/Dataproc pipelines and streaming Beam/Dataflow pipelines with incremental synchronization and exactly-once processing.
- Manage Avro and Protobuf schema versioning, compatibility rules, migrations, and registry standards.
- Set orchestration standards with Flyte and evaluate or benchmark other workflow platforms.
- Design multi-tenant isolation, quotas, contamination validation, and customer-environment data boundaries.
- Operate feature stores, vector databases, embedding upsert pipelines, RAG refresh workflows, and retrieval-context freshness SLAs.
- Build data anonymizer, data labeler, synthetic-data, service API, and data-quality validation capabilities.
- Expose feature-serving, embedding, and schema-validation services over HTTPS, mTLS, and gRPC where appropriate.
Requirements
- At least 1 year of experience as a Principal Software Engineer at a SaaS company.
- Demonstrated organization-wide principal impact through adopted platform standards or major cross-team pipeline and schema migrations.
- Production ownership of an end-to-end data lake, including layout, partitioning, retention tiers, table formats, compaction, and access control.
- Deep Spark experience with PySpark or Scala, including executor tuning, shuffle diagnosis, and Iceberg maintenance.
- Hands-on production experience with Apache Beam and Dataflow, including windowing, exactly-once processing, side inputs, and autoscaling.
- Experience with Protobuf or Avro schema compatibility rules and production breaking-change migrations.
- Production-scale orchestration experience with Flyte, Kubeflow Pipelines, Airflow, or Prefect.
- Experience designing multi-tenant data architectures with strict tenant isolation.
- Experience operating Feast or Tecton feature stores with point-in-time joins and online/offline consistency.
- Production experience with Pgvector or Qdrant, including index tuning, approximate nearest-neighbor search, and embedding upsert pipelines.
- Understanding of RAG data fundamentals, including chunking, embedding model selection, retrieval-quality evaluation, and context freshness.
- Experience with gRPC and HTTPS/mTLS, proto contracts, and certificate lifecycle management.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience or equivalent military experience.
- Preferred: differential privacy or k-anonymity experience, relevant open-source contributions, IAM/access-governance data familiarity, or Iceberg/Delta Lake experience at petabyte scale.
Benefits
- Competitive compensation, benefits, and growth opportunities.
- Opportunity to work on a large-scale Kubernetes-based SaaS platform and solve cloud and reliability problems at scale.
- Collaboration with strong engineers in a reliability-focused culture.
Tech Stack
Categories
Data EngineeringML Engineering
