Saviynt

Principal Software Engineer, AI Platform Engineering

Saviynt
Apply
4 months ago
Milpitas, CA, USAStaff+

Base Salary

$274k - $304k/yr

Responsibilities

  • Define architectural standards for governed, tenant-isolated, PII-free, traceable AI training data.
  • Own the Google Cloud Storage data lake layout, tiering, encryption, lifecycle rules, retention, access control, and Iceberg maintenance.
  • Build and operate batch Spark/Dataproc pipelines and streaming Beam/Dataflow pipelines with incremental synchronization and exactly-once processing.
  • Manage Avro and Protobuf schema versioning, compatibility rules, migrations, and registry standards.
  • Set orchestration standards with Flyte and evaluate or benchmark other workflow platforms.
  • Design multi-tenant isolation, quotas, contamination validation, and customer-environment data boundaries.
  • Operate feature stores, vector databases, embedding upsert pipelines, RAG refresh workflows, and retrieval-context freshness SLAs.
  • Build data anonymizer, data labeler, synthetic-data, service API, and data-quality validation capabilities.
  • Expose feature-serving, embedding, and schema-validation services over HTTPS, mTLS, and gRPC where appropriate.

Requirements

  • At least 1 year of experience as a Principal Software Engineer at a SaaS company.
  • Demonstrated organization-wide principal impact through adopted platform standards or major cross-team pipeline and schema migrations.
  • Production ownership of an end-to-end data lake, including layout, partitioning, retention tiers, table formats, compaction, and access control.
  • Deep Spark experience with PySpark or Scala, including executor tuning, shuffle diagnosis, and Iceberg maintenance.
  • Hands-on production experience with Apache Beam and Dataflow, including windowing, exactly-once processing, side inputs, and autoscaling.
  • Experience with Protobuf or Avro schema compatibility rules and production breaking-change migrations.
  • Production-scale orchestration experience with Flyte, Kubeflow Pipelines, Airflow, or Prefect.
  • Experience designing multi-tenant data architectures with strict tenant isolation.
  • Experience operating Feast or Tecton feature stores with point-in-time joins and online/offline consistency.
  • Production experience with Pgvector or Qdrant, including index tuning, approximate nearest-neighbor search, and embedding upsert pipelines.
  • Understanding of RAG data fundamentals, including chunking, embedding model selection, retrieval-quality evaluation, and context freshness.
  • Experience with gRPC and HTTPS/mTLS, proto contracts, and certificate lifecycle management.
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience or equivalent military experience.
  • Preferred: differential privacy or k-anonymity experience, relevant open-source contributions, IAM/access-governance data familiarity, or Iceberg/Delta Lake experience at petabyte scale.

Benefits

  • Competitive compensation, benefits, and growth opportunities.
  • Opportunity to work on a large-scale Kubernetes-based SaaS platform and solve cloud and reliability problems at scale.
  • Collaboration with strong engineers in a reliability-focused culture.

Tech Stack

Apache AirflowApache BeamApache SparkdbtgRPCKubernetesRedisScala

Categories

Data EngineeringML Engineering
Saviynt

About Saviynt

1,001-5,000 employees
Contact me