
Senior Associate/Assistant Vice President, AI Data Engineer
Temasek Holdings (Private) Limited1 month ago
Singapore, SingaporeMid Level / Senior
Responsibilities
- Design and build AI-ready data architectures using structured data stores, vector knowledge bases, graph databases, and hybrid retrieval systems.
- Build and maintain RAG data layers, including document ingestion, chunking, embedding generation, metadata tagging, and vector index management.
- Define ontology and schema standards for AI-accessible data assets.
- Architect real-time and near-real-time market, research, news, and portfolio data feeds with enforced freshness and latency SLAs.
- Implement data quality standards, automated quality gates, lineage tracking, observability, drift detection, schema alerts, and SLA dashboards.
- Ensure AI data assets meet data classification, access control, governance, audit, and cross-border data handling requirements.
- Build reusable investment knowledge graphs, data APIs, document intelligence pipelines, and portfolio analytics data services.
- Maintain the AI data catalogue and contribute to enterprise data platform strategy.
- Evaluate and manage external data vendors, market data providers, data quality assessments, and licensing agreements.
Requirements
- 4–8 years of data engineering experience, including at least 2 years building production data infrastructure for AI/ML or LLM-powered systems.
- Experience working in data-intensive organizations with complex heterogeneous environments; investment, financial, or enterprise data experience is preferred.
- Production experience building RAG pipelines or AI knowledge bases, including vector store management, embedding pipelines, and chunking strategy optimization.
- Strong fundamentals in pipeline design and orchestration, schema design, data quality frameworks, and lineage tracking.
- Proficiency with Python, pandas, Polars, SQLAlchemy, dbt, Apache Airflow or Prefect, and Spark.
- Experience with batch and streaming architectures using Kafka, Kinesis, or equivalent technologies.
- Experience with vector databases such as Pinecone, Weaviate, pgvector, and Chroma, plus embedding model management.
- Experience with LlamaIndex or LangChain data connectors, document parsing, OCR tooling, and document chunking strategies.
- SQL proficiency across multiple dialects and experience with Snowflake, BigQuery, Redshift, or Databricks.
- Familiarity with graph database concepts and Neo4j or equivalent.
- Experience with data quality tools such as Great Expectations or Soda, lineage tools such as OpenLineage, DataHub, or Marquez, and observability platforms such as Monte Carlo or Acceldata.
- Exposure to investment data domains, including company financials, market data, and portfolio systems, is preferred.
Tech Stack
Amazon RedshiftApache AirflowApache KafkaApache SparkDatabricksdbtGoogle BigQueryNeo4jPandasPythonSnowflakeSQL
Categories
Data Engineering