Tekion

Staff Software Engineer - AI Engineer

Tekion
Apply
3 months ago
Bengaluru, IndiaStaff+

Responsibilities

  • Build and operate the LLM control plane and gateway with routing, quotas, failover, and token and cost tracking.
  • Ship unified APIs and SDKs with normalized schemas, structured outputs, caching, and observability through traces, logs, and metrics.
  • Enforce content filtering, prompt and response validation, privacy protections, and PII redaction.
  • Enable multi-model and multi-vendor LLM usage with automated canarying and versioning.
  • Own the agent runtime, including tool registries, permissions, function calling, grounding, retrieval, orchestration, state, and long-running workflows.
  • Build platform components for classical ML training and scoring pipelines, experiment tracking, packaging, monitoring, drift detection, retraining, and tuning.
  • Evolve the domain graph, entity resolution, ingestion pipelines, and real-time contextual data services with access controls and lineage.
  • Implement hybrid graph, vector, and keyword retrieval with caching and TTL controls.
  • Run offline and online evaluations for quality, factuality, bias, and safety.
  • Define latency, uptime, and cost objectives and enable autoscaling and spend controls.
  • Maintain model and agent registries with versioning, approvals, audit trails, compliance support, and reproducibility.
  • Provide templates, CLIs, sandboxes, and documentation while mentoring engineers and promoting MLOps and AI safety practices.

Requirements

  • 8+ years building large-scale data, machine learning, or platform systems.
  • Strong software engineering fundamentals in API design, concurrency, and distributed systems.
  • Production experience with Python plus one of Java, Scala, or Go.
  • Experience with microservices, API design, and multi-tenant SaaS performance, reliability, and cost engineering.
  • MLOps experience with training pipelines, experiment tracking and registries, model CI/CD, A/B testing, shadowing, canarying, and online feature computation.
  • Experience building or operating an LLM gateway or control plane with provider adapters, routing, policies, caching, quotas, rate limits, and cost and token accounting.
  • Experience with agentic systems, tool use, function calling, orchestration frameworks, human-in-the-loop workflows, safety guardrails, and online evaluation telemetry.
  • Knowledge of practical machine learning, including feature engineering, training, evaluation, drift detection, and deploying models for user-facing workflows.
  • Experience with knowledge graphs, GraphQL, vector search, and hybrid retrieval patterns.
  • Experience with cloud infrastructure and containers, particularly AWS, Docker, and Kubernetes.
  • Ability to mentor engineers, improve developer experience, and establish AI safety and MLOps best practices.

Tech Stack

Apache AirflowApache FlinkApache KafkaApache SparkAWSDockerGoGraphQLgRPCJavaKubernetesLightGBMMLflowNeo4jPythonScalaXGBoost
Tekion

About Tekion

1,001-5,000 employees

Tekion builds an AI-native, cloud platform for automotive retail that unifies dealers, OEMs, and partners. Its products include Automotive Retail Cloud (a dealership management system for retailers), Automotive Enterprise Cloud for manufacturers, and Automotive Partner Cloud for integrations, delivered as subscription software. Privately held and headquartered in Pleasanton, California, Tekion raised private equity funding in 2024.

Contact me