
Distinguished Engineer, Data Platforms
S&P Global1 day ago
Hyderābād, IndiaStaff+
Responsibilities
- Plan and execute the transition of existing data pipelines to the target enterprise Databricks data platform.
- Design repeatable ingestion, transformation, testing, deployment, monitoring, and onboarding patterns for batch and streaming data.
- Guide cloud-native AWS data pipeline architecture using Amazon S3, AWS Glue, AWS Lambda, Amazon Kinesis, AWS Lake Formation, and AWS IAM.
- Influence architecture for Databricks, Delta Lake, Apache Iceberg, Unity Catalog, metadata-driven processing, and governed data lakes.
- Create context engineering practices, context packs, prompt libraries, and AI-enabled playbooks for data platform engineering.
- Use LLM-based tools to support code generation, refactoring, SQL and PySpark optimization, pipeline analysis, documentation, and troubleshooting.
- Integrate platform pipelines with data mastering capabilities and support reconciliation, metadata alignment, stewardship, golden-copy generation, and trusted distribution.
- Support semantic modeling, business glossaries, data catalogs, lineage, governed data products, and consistent business definitions.
- Provide technical direction, architecture guidance, design reviews, code reviews, implementation planning, and mentorship without direct people management.
- Establish standards for infrastructure automation, release automation, environment promotion, observability, reliability, incident management, governance, and responsible AI use.
Requirements
- At least 10 years of experience in data engineering, data platforms, cloud data architecture, software engineering, or related engineering domains.
- Experience operating as a senior individual contributor, technical lead, principal engineer, staff engineer, distinguished engineer, architect, or equivalent in complex data platform environments.
- Proven ability to influence architecture, provide technical guidance, and drive engineering practices across distributed teams.
- Strong hands-on experience with modern batch and/or streaming data engineering and pipeline development.
- Experience with Databricks-based data platforms and pipeline migration to lakehouse-oriented architectures.
- Strong AWS cloud-native data engineering experience, including Amazon S3, AWS Glue, AWS Lambda, AWS Lake Formation, Amazon Kinesis, security, orchestration, and monitoring capabilities.
- Working knowledge of Delta Lake, Apache Iceberg, Databricks Unity Catalog, metadata-driven pipelines, data catalogs, and governed lakehouse architectures.
- Experience with data mastering, MDM, reference data, market data, investment data, risk data, or trusted data distribution workflows.
- Understanding of data quality, schema management, lineage, metadata, access control, encryption, observability, and production support.
- Practical experience using Claude, GitHub Copilot, LLMs, or similar AI-assisted engineering tools, including context engineering and evaluation of generated outputs.
- Ability to work across global and offshore teams and translate architectural direction into actionable designs, implementation patterns, and delivery plans.
- Preferred qualifications include NeoXam DataHub, semantic modeling, enterprise cataloging and metadata management, Databricks or AWS certifications, and equivalent practical platform experience.
Benefits
- Health and wellness coverage.
- Generous flexible time off.
- Continuous learning resources and professional development support.
- Retirement planning, continuing education support, company-matched student loan contribution, and financial wellness programs.
- Family-friendly benefits, retail discounts, and referral incentive awards.
- Hyderabad location with a working model of twice per week or nine days per month in the office, and a stated shift of 12:00 to 21:00 IST.
Tech Stack
DatabricksSQL
Categories
Data Engineering