3 months ago
Bengaluru, IndiaSenior
Responsibilities
- Build and operate the LLM control plane and gateway with routing, quotas, failover, and token and cost tracking.
- Ship unified APIs and SDKs with normalized schemas, structured outputs, caching, and observability through traces, logs, and metrics.
- Implement content filtering, prompt and response validation, PII redaction, permissions, audit trails, and safe-actioning controls.
- Support multi-model and multi-vendor LLM usage with automated canarying, versioning, and provider portability.
- Own the agent runtime, including tool registration, function calling, grounding, retrieval, orchestration, state, and long-running workflows.
- Build training, scoring, experiment-tracking, packaging, monitoring, drift-detection, retraining, and tuning components for classical ML and deep models.
- Evolve the domain graph, entity resolution, data ingestion pipelines, real-time context services, and access-control and lineage capabilities.
- Power hybrid graph, vector, and keyword retrieval with caching and TTL controls.
- Run offline and online evaluations for quality, factuality, bias, and safety.
- Define latency, uptime, and cost SLOs and enable autoscaling and spend controls.
- Maintain model and agent registries with versioning, approvals, reproducibility, and compliance support.
- Provide templates, CLIs, sandboxes, and documentation while mentoring engineers and promoting MLOps and AI safety practices.
Requirements
- 5+ years building large-scale data, ML, or platform systems.
- Strong software engineering fundamentals, including API design, concurrency, and distributed systems.
- Production experience with Python and one of Java, Scala, or Go.
- Experience with microservices, API design, MLOps pipelines, model tracking and registries, model CI/CD, A/B testing, shadowing, canarying, and online feature computation.
- Experience with AWS, Docker, Kubernetes, and performance, reliability, and cost engineering for multi-tenant SaaS.
- Practical ML knowledge covering feature engineering, training, evaluation, drift detection, and deployment of models supporting user-facing workflows.
- Experience operating an LLM gateway or control plane with provider adapters, routing, policies, caching, quotas, rate limits, and token and cost accounting.
- Experience with agentic systems, tool use, function calling, orchestration frameworks, human-in-the-loop workflows, safety guardrails, and online evaluation and telemetry.
- Experience with knowledge graphs, GraphQL, vector search, and hybrid retrieval patterns.
- Preferred mindset includes platform-as-product thinking, strong observability and access-control practices, cost awareness, vendor neutrality, documentation, and mentoring.
Tech Stack
Apache AirflowApache FlinkApache KafkaApache SparkAWSDockerGoGraphQLgRPCJavaKubernetesLightGBMMLflowNeo4jPythonScalaXGBoost
Categories
About Tekion
Tekion builds an AI-native, cloud platform for automotive retail that unifies dealers, OEMs, and partners. Its products include Automotive Retail Cloud (a dealership management system for retailers), Automotive Enterprise Cloud for manufacturers, and Automotive Partner Cloud for integrations, delivered as subscription software. Privately held and headquartered in Pleasanton, California, Tekion raised private equity funding in 2024.
