5 hours ago
Base Salary
$192k - $455k/yr
Responsibilities
- Design and build data pipelines that retrieve, parse, transform, and load new data sources, primarily as AWS Glue jobs written in PySpark.
- Investigate unfamiliar source systems, data formats, semantics, access patterns, and constraints before implementation.
- Design schemas and data models, including analytical models in ClickHouse, with appropriate partitioning and indexing decisions.
- Deploy, operate, package, and support integrations across office-accessible and restricted customer environments.
- Monitor pipeline health, data quality, throughput, and freshness, and build dashboards and alerts using the LGTM stack.
- Diagnose production failures, perform root-cause analysis, lead incident resolution, and improve runbooks and system designs.
- Participate in an on-call rotation using PagerDuty.
- Collaborate with forward deployed engineers and turn field observations into platform capabilities.
Requirements
- 5+ years of experience in data engineering or software engineering with a substantial data infrastructure focus.
- Expert Python skills and experience designing and maintaining production-grade codebases.
- Hands-on experience building and operating ETL/ELT pipelines with Spark and PySpark, ideally as AWS Glue jobs.
- Strong SQL and schema design skills, including partitioning and indexing decisions.
- Experience with column-oriented analytical databases such as ClickHouse, Redshift, or BigQuery.
- Experience debugging production data systems and root-causing issues under time pressure.
- Ability to work independently on production systems and make sound escalation decisions.
- Ability to work full-time on-site in New York City, complete initial onboarding in Arlington, Virginia, and travel there occasionally afterward.
- U.S. citizenship and eligibility to obtain a U.S. Government security clearance.
- Preferred experience with Apache Iceberg, Delta Lake, Trino, Presto, Athena, Kafka, NATS, Kinesis, RabbitMQ, Grafana, Datadog, Splunk, air-gapped or classified environments, Docker, container orchestration, and CI/CD concepts.
Benefits
- Medical, dental, vision, life/AD&D, and disability plan options.
- Paid parental leave, including 12 weeks for birthing parents, 4 weeks for non-birthing parents, and 6 weeks for adoptive, foster, or intended parents through surrogacy.
- Paid holidays and flexible PTO.
- 401(k) with pre-tax and Roth options, HSA/FSA options, and dependent care FSA.
- Full-time on-site work in New York City, with roughly two weeks of initial onboarding in Arlington, Virginia, and occasional travel to Arlington afterward.
Tech Stack
Amazon RedshiftApache KafkaApache SparkAWSClickHouseDatadogDockerGoGoogle BigQueryGrafanaPostgreSQLPrestoPythonRabbitMQSplunkSQLTypeScript
Categories
Data Engineering