20 hours ago
Bengaluru, IndiaStaff+
Responsibilities
- Lead the design and development of hybrid retrieval architectures using vector similarity search and structured graph traversals.
- Architect scalable ingestion, embedding, and indexing pipelines for massive multimodal datasets.
- Develop advanced retrieval capabilities including multi-stage re-ranking, graph tooling for LLMs, and dynamic metadata filtering.
- Design knowledge-graph schemas and semantic data models supporting high-performance relationship mapping and ontological integrity.
- Build automated data validation and drift-detection systems for embedding quality and vector-space health.
- Implement persistent, observable AI-agent memory systems targeting sub-second latency.
- Unify disparate data sources into coherent, searchable knowledge bases and establish data-organization standards.
- Evaluate and integrate emerging database and retrieval technologies with AI Research and Product teams.
Requirements
- 8+ years of experience in data engineering or backend systems focused on high-performance data retrieval and storage.
- BE/B.Tech in Computer Science, Mathematics, or an equivalent qualification; an MS or PhD is a plus.
- Expert proficiency in Python, Java, or Go and strong knowledge of distributed-system design patterns.
- Deep understanding of vector-database indexing strategies including HNSW, IVFFlat, and PQ, plus cosine, Euclidean, and dot-product distance metrics.
- Experience with Pinecone, Milvus, Weaviate, or Qdrant.
- Strong experience with graph databases such as Neo4j, AWS Neptune, or ArangoDB and query languages including Cypher or Gremlin.
- Experience building semantic layers, ontologies, taxonomies, and other data models.
- Hands-on experience with LangChain, LlamaIndex, and embedding models or providers including OpenAI, HuggingFace, and Cohere.
- Experience with large-scale data processing using Spark, Flink, or Kafka for real-time indexing and ETL.
- Understanding of information-retrieval techniques including BM25, TF-IDF, and reciprocal rank fusion.
- Experience with AWS, GCP, or Azure and Kubernetes.
- Preferred: B-Tech or M-Tech in Computer Science or a related technical discipline and 9+ years in Infrastructure, DevOps, or SRE roles focused on production traffic, automation, and API delivery.
Tech Stack
Categories
BackendData Engineering
About Hotstar
We’ve moved! 🌟 Follow us on @JioHotstar for all the latest stories, updates & more.
