7 months ago
Bengaluru, IndiaSenior
Responsibilities
- Build and operate the LLM control plane and gateway with routing, quotas, failover, and token and cost tracking.
- Ship unified REST/gRPC APIs and SDKs with normalized schemas, structured outputs, caching, and tracing, logging, and metrics.
- Enforce content filtering, prompt and response validation, PII redaction, privacy, and safety controls.
- Enable multi-model and multi-vendor LLM usage with automated canarying and versioning.
- Own the agent runtime, including tool registration, permissions, function calling, grounding, retrieval, orchestration, state, and long-running workflows.
- Build training and scoring platform components for classical and deep ML models and standardize experiment tracking and packaging.
- Monitor model and data drift and support retraining and tuning to maintain model accuracy and relevance.
- Add human review and safe actioning before agents interact with dealer systems.
- Evolve the domain graph and entity resolution systems and build reliable data ingestion pipelines.
- Serve real-time contextual data with access controls and lineage and support hybrid graph, vector, and keyword retrieval.
- Run offline and online evaluations for quality, factuality, bias, and safety.
- Define latency, uptime, and cost SLOs and enable autoscaling and spend controls.
- Maintain model and agent registries, versioning, approvals, audit trails, compliance support, and reproducibility.
- Provide templates, CLIs, sandboxes, and documentation for product teams while mentoring engineers and promoting MLOps and AI safety practices.
Requirements
- 5+ years building large-scale data, machine learning, or platform systems.
- Strong software engineering fundamentals in API design, concurrency, and distributed systems.
- Production experience with Python plus one of Java, Scala, or Go.
- Experience with microservices, API design, MLOps pipelines, model tracking and registries, model CI/CD, A/B testing, shadowing, canarying, and online feature computation.
- Experience with AWS, Docker, Kubernetes, and performance, reliability, and cost engineering for multi-tenant SaaS.
- Practical machine learning knowledge covering feature engineering, training, evaluation, drift detection, and deployment of models in user-facing workflows.
- Experience building or operating an LLM gateway or control plane with provider adapters, routing, policies, caching, quotas, rate limits, and cost and token accounting.
- Experience with agentic systems, tool use, function calling, orchestration, human-in-the-loop workflows, safety guardrails, and online evaluation and telemetry.
- Experience with knowledge graphs, GraphQL, vector search, and hybrid retrieval patterns.
- Ability to design developer-focused platform products, observability, access controls, graceful fallbacks, documentation, and teaching.
Tech Stack
Apache AirflowApache FlinkApache KafkaApache SparkAWSDockerGoGraphQLgRPCJavaKubernetesLightGBMMLflowNeo4jPythonScalaXGBoost
Categories
About Tekion
Tekion builds an AI-native, cloud platform for automotive retail that unifies dealers, OEMs, and partners. Its products include Automotive Retail Cloud (a dealership management system for retailers), Automotive Enterprise Cloud for manufacturers, and Automotive Partner Cloud for integrations, delivered as subscription software. Privately held and headquartered in Pleasanton, California, Tekion raised private equity funding in 2024.
