S&P Global

Distinguished Engineer, Data Platforms

S&P Global
Apply
1 day ago
Hyderābād, IndiaStaff+

Responsibilities

  • Plan and execute the transition of existing data pipelines to the target enterprise Databricks data platform.
  • Design repeatable ingestion, transformation, testing, deployment, monitoring, and onboarding patterns for batch and streaming data.
  • Guide cloud-native AWS data pipeline architecture using Amazon S3, AWS Glue, AWS Lambda, Amazon Kinesis, AWS Lake Formation, and AWS IAM.
  • Influence architecture for Databricks, Delta Lake, Apache Iceberg, Unity Catalog, metadata-driven processing, and governed data lakes.
  • Create context engineering practices, context packs, prompt libraries, and AI-enabled playbooks for data platform engineering.
  • Use LLM-based tools to support code generation, refactoring, SQL and PySpark optimization, pipeline analysis, documentation, and troubleshooting.
  • Integrate platform pipelines with data mastering capabilities and support reconciliation, metadata alignment, stewardship, golden-copy generation, and trusted distribution.
  • Support semantic modeling, business glossaries, data catalogs, lineage, governed data products, and consistent business definitions.
  • Provide technical direction, architecture guidance, design reviews, code reviews, implementation planning, and mentorship without direct people management.
  • Establish standards for infrastructure automation, release automation, environment promotion, observability, reliability, incident management, governance, and responsible AI use.

Requirements

  • At least 10 years of experience in data engineering, data platforms, cloud data architecture, software engineering, or related engineering domains.
  • Experience operating as a senior individual contributor, technical lead, principal engineer, staff engineer, distinguished engineer, architect, or equivalent in complex data platform environments.
  • Proven ability to influence architecture, provide technical guidance, and drive engineering practices across distributed teams.
  • Strong hands-on experience with modern batch and/or streaming data engineering and pipeline development.
  • Experience with Databricks-based data platforms and pipeline migration to lakehouse-oriented architectures.
  • Strong AWS cloud-native data engineering experience, including Amazon S3, AWS Glue, AWS Lambda, AWS Lake Formation, Amazon Kinesis, security, orchestration, and monitoring capabilities.
  • Working knowledge of Delta Lake, Apache Iceberg, Databricks Unity Catalog, metadata-driven pipelines, data catalogs, and governed lakehouse architectures.
  • Experience with data mastering, MDM, reference data, market data, investment data, risk data, or trusted data distribution workflows.
  • Understanding of data quality, schema management, lineage, metadata, access control, encryption, observability, and production support.
  • Practical experience using Claude, GitHub Copilot, LLMs, or similar AI-assisted engineering tools, including context engineering and evaluation of generated outputs.
  • Ability to work across global and offshore teams and translate architectural direction into actionable designs, implementation patterns, and delivery plans.
  • Preferred qualifications include NeoXam DataHub, semantic modeling, enterprise cataloging and metadata management, Databricks or AWS certifications, and equivalent practical platform experience.

Benefits

  • Health and wellness coverage.
  • Generous flexible time off.
  • Continuous learning resources and professional development support.
  • Retirement planning, continuing education support, company-matched student loan contribution, and financial wellness programs.
  • Family-friendly benefits, retail discounts, and referral incentive awards.
  • Hyderabad location with a working model of twice per week or nine days per month in the office, and a stated shift of 12:00 to 21:00 IST.

Tech Stack

DatabricksSQL

Categories

Data Engineering
Contact me