5 months ago
Bellevue, WA, USAMid Level
Responsibilities
- Build and maintain data pipeline lifecycles, production transformation logic, operational outputs, schema contracts, and measurable validation.
- Use AI-assisted workflows for data exploration, transformation, querying, investigation, refactoring, schema discovery, SQL generation, and troubleshooting.
- Build integration points between data pipelines and ML pipelines, including schema-bound datasets, ML-ready inputs, and semantic-layer outputs.
- Operate production systems through monitoring, alerting, incident remediation, automated checks, regression coverage, and durable bug fixes.
- Contribute to schema design and enforce data contracts that separate model logic from the system of record.
- Partner with product, science, and platform teammates to clarify requirements, identify tradeoffs, and deliver customer-ready work.
Requirements
- Degree in Computer Science, Mathematics, Statistics, or another data-intensive discipline, or equivalent practical experience.
- 4+ years of professional development experience with strong hands-on SQL and Python in production.
- 3+ years of experience with structured and semi-structured data, modern warehouses or lakehouses, and schema design in evolving domains.
- Experience with lakehouse or warehouse patterns, incremental processing, notebook workflows, and basic performance and cost awareness.
- Experience with production-system ownership, debugging, reliability improvement, data validation, data quality checks, and regression protection.
- Strong communication, collaboration, follow-through, and judgment when verifying AI-generated SQL or pipelines.
- Spark or equivalent large-scale batch processing experience is preferred; Scala, Flink, or Beam experience is a plus.
- Experience in supply chain, planning, or fulfillment domains is a plus.
Tech Stack
Categories
Data Engineering
