Citi

Big data/Python/Databricks Engineer Engineer

Citi
Apply
1 day ago
Pune, IndiaSenior

Responsibilities

  • Build and maintain scalable Python and PySpark data pipelines for structured and unstructured data across distributed platforms.
  • Develop and optimize data workflows using Hadoop, HDFS, and Hive to support downstream reporting and analytics.
  • Manage the data application development lifecycle from analysis and design through testing, implementation, and production support.
  • Conduct feasibility studies, time and cost estimates, and technical planning for data engineering delivery.
  • Collaborate with analytics and reporting teams to build and support Tableau dashboards and translate data into business insights.
  • Administer and troubleshoot data processes in Linux environments, including operational stability and performance support.
  • Assess engineering risks and align solutions with security, data governance, and compliance standards.
  • Support data engineering projects, delivery planning, deadlines, and unexpected changes in requirements.

Requirements

  • 5 to 8 years of experience in data engineering, big data development, or software application development focused on large-scale data platforms.
  • Hands-on experience with Python and PySpark for large-scale data ingestion, transformation, and pipeline orchestration.
  • Expertise in ETL, Oracle DB, and SQL, with knowledge of Databricks.
  • Practical knowledge of Hadoop ecosystem technologies including HDFS, Hive, Hadoop cluster operations, and Ozone storage.
  • Experience working in Linux environments with scripting, job scheduling, and process management.
  • Experience building or supporting Tableau dashboards and reports.
  • Familiarity with RStudio for statistical analysis or data exploration.
  • Deep expertise in large language models, including OpenAI, Gemini, Claude, Llama, and local or open-source models.
  • Bachelor's degree or equivalent experience in a relevant technical discipline.
  • Preferred experience with cloud-native data platforms, migration from on-premise Hadoop environments, data governance, metadata management, data quality frameworks, and technology project management techniques.

Benefits

  • Hybrid working model with 2 days in the office and 3 days working remotely.
  • Continuous learning and development resources across data engineering, analytics, and cloud technologies.
  • Opportunity to work on large-scale enterprise data engineering challenges.
  • Technical autonomy, limited direct supervision, and the opportunity to act as a subject matter expert for senior stakeholders.
  • Structured career progression within Citi's global technology organization.
  • Competitive financial wellbeing and benefits package aligned to location and level.
  • Full-time employment.

Tech Stack

Apache HadoopApache HiveDatabricksLinuxOracle DatabasePythonSQL

Categories

Data Engineering
Citi

About Citi

10,000+ employees

Citi is a public financial-services company offering consumer and institutional banking, credit cards, wealth management, treasury and trade solutions, and capital-markets services. It serves individuals, corporations, financial institutions, and governments in more than 160 countries and jurisdictions, earning interest and fee income from lending, payments, trading, and advisory. Founded in 1812 and headquartered in New York, it trades on the NYSE under the ticker C.

Contact me