Citi

Big Data PySpark Lead Engineer - Vice President

Citi
Apply
1 month ago
Jersey City, NJ, USAStaff+
H1B sponsor

Base Salary

$142k - $213k/yr

Responsibilities

  • Build and maintain scalable PySpark data pipelines for large volumes of structured and unstructured data.
  • Design and develop data solutions across Hadoop, Hive, HDFS, Sqoop, Spark, Impala, and Scala.
  • Develop and manage real-time and batch workflows with high availability and low-latency delivery.
  • Write complex SQL queries for data extraction, validation, transformation, and analysis across distributed systems.
  • Design scalable data models and architecture patterns aligned with data warehouse and dimensional modeling principles.
  • Automate pipeline scheduling and orchestration using shell scripting and Autosys or equivalent tools.
  • Identify, assess, and resolve technical risks and data issues while maintaining data platform integrity.
  • Lead application systems analysis and programming activities, define system enhancements, and ensure alignment with architecture standards.
  • Advise and coach mid-level developers and analysts and allocate work as needed.
  • Partner with management teams and communicate technical concepts to technical and non-technical audiences.

Requirements

  • 6–10 years of relevant experience in applications development or systems analysis.
  • Hands-on expertise in PySpark and Big Data processing, including distributed workflow optimization at scale.
  • Production experience with the Hadoop ecosystem, including Hive, HDFS, Sqoop, Spark, Impala, and Scala.
  • Proficiency in complex SQL development for analysis, transformation, and validation of large datasets.
  • Strong understanding of distributed systems architecture, data modeling, data design, data warehouses, and dimensional modeling.
  • Experience with shell scripting and Autosys or equivalent workflow automation tools.
  • Strong analytical, problem-solving, communication, leadership, and project management skills.
  • Bachelor’s or university degree, or equivalent experience; a master’s degree is preferred.
  • Familiarity with Apache Kafka or equivalent streaming technologies, cloud-based Big Data environments, modern data lakes, financial services, or regulated industries is beneficial.

Benefits

  • Full-time position based in Jersey City, New Jersey.
  • Medical, dental, and vision coverage.
  • 401(k), life, accident, and disability insurance.
  • Wellness programs, vacation, sick leave, and paid holidays.
  • Eligible employees may receive discretionary and formulaic incentive and retention awards.

Tech Stack

Apache HadoopApache HiveApache KafkaApache SparkScalaSQL

Categories

Data Engineering
Citi

About Citi

10,000+ employees

Citi's mission is to serve as a trusted partner to our clients by responsibly providing financial services that enable growth and economic progress. Our core activities are safeguarding assets, lending money, making payments and accessing the capital markets on behalf of our clients. We have over 200 years of experience helping our clients meet the world's toughest challenges and embrace its greatest opportunities. We are Citi, the global bank – an institution connecting millions of people across hundreds of countries and cities. For information on Citi’s commitment to privacy, visit on.citi/privacy.

Contact me