
Senior AI Engineer — Inference & Agent Systems
Arcana Analytics7 months ago
Bengaluru, IndiaSenior
Responsibilities
- Optimize inference latency, including achieving sub-400ms time to first token, streaming partial output, KV caching, prompt compression, and dynamic context management.
- Build multi-provider model routing across OpenAI, Anthropic, Gemini, and open-weight models based on latency, cost, and task type.
- Design and implement parallel Plan-Execute-Synthesize agent pipelines using DAGs and build reliable Temporal orchestration with retries, timeouts, recovery, and idempotency.
- Implement structured output validation, retry handling for malformed LLM output, graceful degradation, and reliable tool-call schemas across providers.
- Own the evaluation harness, including ground-truth datasets, automated scoring, LLM-as-judge pipelines, regression detection, latency testing, and adversarial test cases.
- Optimize model serving, cold starts, and asynchronous workers while providing observability across tokens, tool calls, and synthesis steps.
Requirements
- Production experience shipping systems at meaningful scale with a deep understanding of performance and reliability.
- Strong experience with inference pipelines, multi-step agent systems, evaluation harnesses, streaming LLM responses, or resilient handling of LLM non-determinism.
- Experience with Go, Python, Temporal, Kafka, PostgreSQL, and Docker is valuable, though depth matters more than exact stack matching.
- Experience with real-time AI products, model-serving infrastructure, agent platforms, or trusted evaluation infrastructure is sought.
- Fine-tuning models, using LangChain or LlamaIndex, or having a strong ML research background without systems exposure are considered weaker but acceptable signals.
Benefits
- Remote work in India (BLR/Remote India)
- High ownership on a small team with decisions shipping to production
- No cover letter required
Tech Stack
Categories
About Arcana Analytics
Arcana Analytics builds a portfolio intelligence platform for hedge funds and asset managers, combining portfolio data with analytics on performance attribution, factor risk, crowding, and ownership using proprietary datasets. It sells enterprise software and data subscriptions used for portfolio construction, real-time risk, and research workflows by institutional investors in global markets, including APAC. Founded in 2022 and headquartered in New York, the company is privately held and raised a Series B in 2026.