
Knowledge Graph Engineer
BNY Mellon18 days ago
Pittsburgh, PA, USASenior / Staff+
H1B sponsor
Responsibilities
- Perform entity resolution and map source data to IDS entities, relationships, and attributes for enterprise knowledge graph integration.
- Design and build scalable batch, streaming, and near-real-time pipelines for internal and external data sources.
- Transform APIs, flat files, streaming data, and unstructured content into standardized datasets aligned to IDS models.
- Develop reusable frameworks for identifier, symbology, unit, hierarchy, and event-data normalization.
- Implement data quality, schema validation, anomaly detection, lineage, provenance, and traceability controls.
- Support multi-vendor ingestion, comparison, reconciliation, source prioritization, and coverage and quality analytics.
- Build modular, cloud-native pipelines optimized for scalability, performance, reliability, and cost efficiency.
- Collaborate with platform, product, data, and ontology teams to deliver production-ready solutions and downstream APIs and data products.
Requirements
- Bachelor's degree in a related discipline or equivalent work experience.
- Typically 8-12 years of experience, with at least 4 years focused on data analysis and business intelligence preferred.
- Expertise with RDF, OWL, SHACL, SPARQL, LPG, Cypher, and GQL, including round-tripping between graph models without semantic drift.
- Experience building ontology-based knowledge graphs and evaluating graph persistence and virtualization options.
- Experience with graph modularization, versioning, temporality, data entitlements, licensing, and usage tracking.
- Understanding of AI-to-knowledge-graph integration patterns, including MCP, Graph RAG, identity resolution, entity and relationship extraction, hybrid KG/vector retrieval, text-to-query, and agentic workflows.
- Experience with production-grade data engineering and high-volume resilient pipelines using Python, Spark, SQL, ETL/ELT frameworks, and orchestration tools.
- Experience designing transformation and normalization layers, schema evolution, backward compatibility, performance tuning, cost optimization, monitoring, logging, and lineage frameworks.
- Expertise with Snowflake, AWS, Databricks, lakehouse architectures, and API-based data integration.
- Financial dataset experience covering market data, pricing, reference data, portfolio holdings, transactions, or corporate actions; familiarity with Bloomberg, ICE, or MSCI is preferred.
- Advanced degree, preferably involving statistics or statistical analysis, is preferred.
Benefits
- The role is located in Pittsburgh, Pennsylvania or Lake Mary, Florida.
Tech Stack
Categories
Data Engineering