4 months ago
Base Salary
$192k - $455k/yr
Responsibilities
- Design and build data pipelines that retrieve, parse, transform, and load new data sources into applications using AWS Glue jobs and PySpark.
- Investigate undocumented source systems and data formats, then design schemas and data models suited to platform performance and data shape.
- Deploy, operate, package, and support integrations in office-accessible, restricted, and customer environments.
- Monitor pipeline health, data quality, throughput, and freshness using the LGTM observability stack, and build dashboards and alerts.
- Diagnose production data issues, resolve incidents, conduct root-cause analysis and post-incident reviews, and create resulting runbooks or fixes.
- Participate in PagerDuty on-call and collaborate with forward deployed engineers and customer deployment teams.
Requirements
- 5+ years of experience in data engineering or software engineering with a substantial data infrastructure focus.
- Expert Python skills and experience designing and maintaining production-grade codebases.
- Hands-on experience building and operating ETL/ELT pipelines with Spark and PySpark, ideally as AWS Glue jobs.
- Strong SQL and schema design skills, including partitioning and indexing decisions based on access patterns.
- Experience with column-oriented analytical databases such as ClickHouse, Redshift, or BigQuery.
- Experience debugging production data systems and root-causing issues under time pressure.
- Ability to work independently on production systems and make sound escalation decisions.
- Ability to work full-time onsite at the New York City office, complete initial onboarding in Arlington, Virginia, and travel there occasionally afterward.
- U.S. citizenship and eligibility to obtain a U.S. Government security clearance are required; an active clearance is not required to start.
- Preferred experience includes Apache Iceberg, Delta Lake, Trino, Presto, Athena, Kafka, NATS, Kinesis, RabbitMQ, Grafana, Datadog, Splunk, restricted environments, Docker, container orchestration, and CI/CD concepts.
Benefits
- Medical, dental, vision, life/AD&D, and disability plan options.
- Paid parental leave, including 12 weeks for birthing parents, 4 weeks for non-birthing parents, and 6 weeks for adoptive, foster, or intended parents through surrogacy.
- Paid holidays and flexible PTO.
- 401(k) with pre-tax and Roth options, HSA/FSA options, and dependent care FSA.
- Full-time onsite work in New York City, with roughly two weeks of onboarding in Arlington, Virginia, and occasional travel to Arlington afterward.
Tech Stack
Amazon RedshiftApache KafkaApache SparkClickHouseDatadogDockerGoGoogle BigQueryGrafanaPostgreSQLPrestoPythonRabbitMQSplunkSQLTypeScript
Categories
Data Engineering