
ETL Java AI Lead Engineer
JPMorgan Chase6 hours ago
Responsibilities
- Design and deliver large-scale ETL/ELT data pipelines for ingestion, transformation, validation, reconciliation, and publishing.
- Architect and operationalize data lake patterns with partitioning, data quality controls, lineage, governance, and reusable data products.
- Build and optimize distributed processing workloads using Apache Spark and formats such as Parquet and Avro.
- Tune compute, storage, Spark workloads, joins, caching, shuffles, skew handling, and data warehouse performance.
- Build and maintain Java services, ingestion components, orchestration helpers, APIs, and data access layers for data workflows.
- Develop integration components and APIs for downstream consumption and platform interoperability.
- Establish code quality, automated testing, observability, security-by-design, and operational readiness practices.
- Lead adoption of enterprise-authorized AI-assisted engineering tools and define standards for validating AI-generated outputs.
- Coach engineers on responsible AI use, secure handling of data, and compliant engineering workflows.
- Lead technical discussions and drive delivery across cross-functional teams.
Requirements
- Formal training or certification in software engineering concepts and 5+ years of applied experience.
- 10+ years of software engineering experience with strong depth in data engineering and big data platforms.
- Strong hands-on Java development experience and proficiency in at least one data-focused language such as Python.
- Experience designing and building robust ETL/ELT pipelines and data integration frameworks.
- Strong Apache Spark and distributed processing experience, including fault tolerance, partitioning, and performance tuning.
- Strong understanding of Parquet, Avro, data lake or lakehouse patterns, data quality, metadata, and governance.
- Experience with Snowflake or an equivalent cloud data warehouse.
- Experience leading AI-assisted software development practices and validating outputs for correctness, performance, and security.
- Understanding of responsible AI use, data sensitivity, secure input and output handling, resiliency, and security expectations.
- Ability to lead technical discussions, communicate with varied stakeholders, and drive cross-functional delivery.
- Preferred experience with Airflow, Dagster, Control-M, Kafka, Spark Structured Streaming, CDC, upserts, Delta Lake, Apache Iceberg, Hudi, production operations, regulated environments, FinOps, and agentic or generative AI solutions.
Benefits
- Competitive total rewards package with eligibility-based benefits and potential incentive compensation.
- Comprehensive health care coverage, on-site health and wellness centers, retirement savings plan, backup childcare, tuition reimbursement, mental health support, and financial coaching.
Tech Stack
Categories
BackendData Engineering
About JPMorgan Chase
JPMorgan Chase provides consumer and commercial banking, payments, credit card, wealth management, and corporate and investment banking services to individuals, businesses, institutions, and governments. The public company (NYSE: JPM) earns revenue from interest, fees, trading, and asset management across operations in more than 100 markets. Headquartered in New York City with roots dating to 1799, it serves retail customers and prominent corporate and government clients through brands including Chase and J.P. Morgan.