about 1 month ago
Responsibilities
- Design and build scalable data platforms, data warehouses, and lakehouse architectures.
- Design and evolve a multi-tenant, cloud-scale data platform serving dealerships and business users.
- Develop data models using dimensional modeling, Data Vault, and enterprise data architecture principles.
- Build and optimize batch, streaming, and CDC-based data ingestion pipelines.
- Help modernize data processing from batch-oriented architectures to near real-time and real-time ingestion.
- Apply query optimization, performance tuning, partitioning, and schema evolution practices.
- Use workflow orchestration, data quality frameworks, and monitoring solutions to support reliable data operations.
- Maintain strong tenant isolation, governance, and data quality across the platform.
Requirements
- 6+ years of experience in data engineering.
- Strong expertise in Python, SQL, and Apache Spark.
- Experience building scalable batch and real-time ETL/ELT pipelines.
- Hands-on experience with AWS services including EMR, S3, Glue, and Athena.
- Experience with Kafka, Flink, or Kinesis for streaming data processing.
- Strong knowledge of dimensional modeling, Data Vault, and data warehousing concepts.
- Experience with Delta Lake, Apache Iceberg, or Apache Hudi.
- Expertise in workflow orchestration using Airflow.
- Experience implementing data quality frameworks and monitoring solutions.
- Strong understanding of partitioning, schema evolution, and performance optimization.
- Familiarity with Git and Infrastructure as Code tools is a plus.
Benefits
- Competitive compensation is offered.
- Generous stock options are available.
- Medical insurance coverage is provided.
- Employees work with experienced technology professionals from leading Silicon Valley companies.
Tech Stack
Categories
Data Engineering