3 months ago
Hyderābād, IndiaMid Level
Responsibilities
- Design and maintain data ingestion and transformation pipelines for LLM training, fine-tuning, and Retrieval-Augmented Generation.
- Build state-managed AI agents and cyclical workflows with LangGraph.
- Architect retrieval layers using document embeddings and semantic search.
- Implement and optimize vector databases, including Pinecone, Weaviate, or Milvus.
- Create scalable schemas for structured and unstructured data supporting AI services.
- Identify and resolve latency bottlenecks in data retrieval and embedding generation.
- Collaborate with AI Researchers and Product Managers to move AI prototypes into production-ready data products.
Requirements
- At least 2 years of professional experience in data engineering or backend-heavy software engineering.
- Expert-level Python skills, including clean, scalable, and asynchronous code.
- Hands-on experience with LangChain or LangGraph for multi-step chains and agentic systems.
- Experience implementing and tuning vector databases for high-volume RAG pipelines.
- Strong understanding of data modeling, ETL/ELT processes, and SQL and NoSQL databases.
- Understanding of embedding models, tokenization, and modern information retrieval techniques.
Categories
AI ApplicationsData Engineering
