1 day ago
Remote, United StatesStaff+
Base Salary
$125k - $182k/yr
Responsibilities
- Lead the design and implementation of advanced data pipelines, ingestion frameworks, and real-time streaming solutions.
- Build, tune, and maintain scalable Kafka producers and consumers using Python, Java, Go, or Scala for sub-second latency and high availability.
- Configure and manage CDC solutions across relational and NoSQL systems and route changes into Kafka topics.
- Develop stream-processing logic to enrich, filter, and aggregate event payloads in flight.
- Implement schema evolution, delivery guarantees, idempotency, data contracts, governance, security, and data quality standards.
- Integrate automated code generation, validation agents, and LLM-driven orchestration for schema drift, unstructured event parsing, and dynamic routing.
- Instrument pipelines with monitoring and alerting, and manage offsets, partition rebalancing, and dead-letter queues.
- Lead technical initiatives, translate complex business requirements into scalable solutions, mentor engineers, conduct code reviews, and guide architectural decisions.
- Drive improvements in automation, efficiency, platform capabilities, resilience, and high availability.
Requirements
- Bachelor’s degree in computer science or a related field, or an equivalent combination of training and experience.
- 6-8+ years of data engineering experience with enterprise-scale platforms.
- 6-8+ years of hands-on experience with Apache Kafka, Confluent Cloud, AWS Kinesis, or Apache Pulsar.
- Advanced proficiency in Python, PySpark, and SQL.
- Production experience with Java, Scala, or Go and Flink, Kafka Streams, or Spark Structured Streaming.
- Strong AWS and modern data platform experience, including CDC tools, Kafka Connect, and log-based replication mechanisms.
- Experience with Apache Iceberg, Delta Lake, Snowflake, or BigQuery for streaming data storage and warehousing.
- Experience with Docker, Kubernetes, Terraform, and CI/CD automation for event-driven microservices.
- Experience implementing data contracts and schema management, including schema evolution and interface reliability.
- Deep understanding of data modeling, ETL/ELT processes, data warehousing, and enterprise data governance.
- Experience leading technical initiatives and mentoring engineers, with strong problem-solving and communication skills.
- Preferred experience includes AI agents for pipeline recovery and schema drift, resilient fallback strategies, event sourcing, CQRS, domain-driven event design, and enterprise event-contract governance.
Benefits
- Medical, dental, vision, and life insurance.
- 401(k) retirement plan with company matching contributions up to 6%, potential discretionary contributions, financial advisory services, and investment options.
- Tuition reimbursement up to $5,250 per year.
- Business-casual environment with the option to wear jeans.
- Paid time off upon hire, including ten paid company holidays and three floating holidays annually.
- 16 hours of paid volunteer time per calendar year.
- Paid parental leave, paid short- and long-term disability, and Family and Medical Leave programs.
- Business Resource Groups supporting inclusion and collaboration.
- Flexible work environment; remote or hybrid employees must provide reliable wired high-speed internet and a suitable home workspace, and office attendance may be required if conditions are inadequate.
- Other necessary computer equipment will be provided.
Tech Stack
Apache FlinkApache KafkaAWSDatadogDockerGoGoogle BigQueryGrafanaJavaKubernetesMongoDBMySQLPostgreSQLPrometheusPythonScalaSnowflakeSQLTerraform
Categories
Data Engineering
