3 months ago
Bengaluru, IndiaStaff+
Responsibilities
- Lead the design and development of hybrid retrieval architectures combining vector similarity search and structured graph traversals.
- Architect scalable pipelines for data ingestion, embedding, indexing, validation, and drift detection across massive multimodal datasets.
- Develop advanced retrieval capabilities including multi-stage re-ranking, graph tooling for LLMs, and dynamic metadata filtering.
- Design knowledge-graph schemas and semantic layers that support high-performance relationship mapping and ontological integrity.
- Build persistent memory systems for AI agents with observability and sub-second latency.
- Unify disparate data sources into coherent, searchable knowledge bases and champion data organization standards.
- Collaborate with AI Research and Product teams to evaluate and integrate emerging database technologies such as HNSW optimizations and GraphRAG.
Requirements
- 8+ years of experience in data engineering or backend systems focused on high-performance data retrieval and storage.
- Bachelor-level degree such as a BE or B.Tech in Computer Science, Mathematics, or a related technical field; an MS or PhD is a plus.
- Expert proficiency in Python, Java, or Go and strong understanding of distributed-system design patterns.
- Deep knowledge of vector databases, indexing strategies including HNSW, IVFFlat, and PQ, and distance metrics including cosine, Euclidean, and dot product.
- Experience with Pinecone, Milvus, Weaviate, or Qdrant.
- Strong background with graph databases such as Neo4j, AWS Neptune, or ArangoDB and query languages such as Cypher or Gremlin.
- Experience building semantic layers, ontologies, taxonomies, and other data models.
- Hands-on experience with LangChain, LlamaIndex, and embedding models or platforms including OpenAI, HuggingFace, and Cohere.
- Proficiency with large-scale data processing using Spark, Flink, or Kafka for real-time indexing and ETL.
- Understanding of information-retrieval fundamentals including BM25, TF-IDF, and reciprocal rank fusion.
- Experience with AWS, GCP, or Azure and Kubernetes.
- Preferred qualifications include a B-Tech or M-Tech from a reputed university and 9+ years in Infrastructure, DevOps, or SRE roles focused on production traffic, automation, and API delivery.
Tech Stack
Categories
BackendData Engineering
About Hotstar
We’ve moved! 🌟 Follow us on @JioHotstar for all the latest stories, updates & more.
