29 days ago
Singapore, SingaporeMid Level
Responsibilities
- Design and develop enterprise LLM applications such as intelligent Q&A systems, knowledge-base assistants, AI copilots, document analysis tools, and customer service agents.
- Build end-to-end RAG systems covering document parsing, chunking, embeddings, hybrid retrieval, reranking, and source-attributed response synthesis.
- Develop prompt engineering, context management, structured output, prompt versioning, and reliability strategies for LLM applications.
- Design AI agents using ReAct, Plan-and-Execute, Reflection, function calling, tool use, and multi-agent collaboration patterns.
- Deploy and optimize open-source LLMs in on-premise and private cloud environments, including model serving, quantization, batching, speculative decoding, and tensor parallelism.
- Build model-serving infrastructure, GPU scheduling, AI gateways, fine-tuning pipelines, domain datasets, evaluation benchmarks, and LLM quality-monitoring systems.
- Implement production monitoring, A/B testing, cost and performance tracking, drift detection, security guardrails, PII detection, prompt-injection defense, and content moderation.
- Track frontier AI research and contribute to internal technical documentation, talks, and best-practice guides.
Requirements
- Bachelor's degree or above in Computer Science, Artificial Intelligence, Machine Learning, or a related technical field; a master's degree or PhD is preferred.
- At least 2 years of professional AI/ML engineering experience, including production deployment of LLM systems at scale.
- Deep understanding of Transformer architecture, attention mechanisms, and LLM pre-training, fine-tuning, and inference.
- Expertise with production LLM application frameworks such as LangChain, LlamaIndex, or Haystack.
- Hands-on experience with RAG, vector databases, embedding models, reranking, hybrid search, query expansion, and HyDE.
- Experience deploying and tuning DeepSeek, Qwen, Kimi, Llama, Mistral, Mixtral, Gemma, or Phi models.
- Strong experience with model-serving infrastructure, GPU scheduling, quantization, inference optimization, KV-cache optimization, and memory-efficient attention.
- Strong Python programming skills and experience with PyTorch, TensorFlow, or JAX; familiarity with FastAPI or Flask for LLM API services.
- Experience with LLM evaluation, A/B testing, and production monitoring of AI systems.
- Preferred experience includes LangGraph, AutoGen, CrewAI, OpenAI Assistants API, MCP, multimodal LLMs, DSPy, PromptLayer, LangSmith, model distillation, cloud GPU providers, open-source AI contributions, or NLP/LLM publications.
