
Member of Technical Staff (Software Engineer, Data Platform)
Perplexity3 months ago
Palo Alto, CA, USA +2 moreSenior / Staff+
H1B Sponsor
Base Salary
$220k - $405k/yr
Responsibilities
- Design and operate large-scale batch and streaming data pipelines for product features, AI training and evaluation, analytics, and experimentation.
- Build event-driven and streaming systems for real-time ingestion, transformation, and delivery, alongside batch frameworks for backfills, aggregations, and offline computation.
- Lead data orchestration architecture, including scheduling, dependency management, retries, service-level agreements, and end-to-end observability.
- Design guarantees for data correctness, freshness, lineage, and recoverability across scale growth, partial failures, and evolving schemas.
- Build self-serve data platforms that enable engineers, data scientists, and analysts to discover data and operate pipelines.
- Improve developer experience through abstractions, paved paths, and standards for data modeling, testing, validation, and deployment.
- Drive architectural decisions across storage, compute, orchestration, and data APIs in partnership with product engineering and data science.
- Mentor engineers, review designs, document systems, and raise the technical bar for data infrastructure.
Requirements
- At least 5 years of software engineering experience for Senior or 8 years for Staff.
- Strong experience building production data infrastructure systems and processing batch or streaming data at scale.
- Deep familiarity with data orchestration systems such as Airflow or Dagster.
- Proficiency in Python and at least one additional backend language such as Go or TypeScript.
- Strong systems thinking regarding reliability, latency, cost, and complexity tradeoffs.
- Experience supporting ML/AI workflows, training pipelines, or evaluation systems.
- Familiarity with data quality, lineage, observability, and governance tooling.
- Prior ownership of internal platforms used by many teams.
Tech Stack
Apache AirflowApache FlinkApache KafkaApache SparkClickHouseDatabricksdbtGoPythonSnowflakeTypeScript
Categories
Data Engineering
About Perplexity
The most powerful answer engine. Powering curiosity with answers backed by up-to-date sources. This is where knowledge begins.