GrepJob
Tekion

Senior Software Engineer - AI Platform

Tekion
Apply
about 1 month ago
Bengaluru, IndiaSenior
H1B Sponsor

Responsibilities

  • Build and operate the LLM control plane and gateway with routing, quotas, failover, and token and cost tracking.
  • Ship unified APIs and SDKs with normalized schemas, structured outputs, caching, and observability across traces, logs, and metrics.
  • Enforce content filtering, prompt and response validation, privacy controls, and PII redaction by default.
  • Enable multi-model and multi-vendor LLM usage with automated canarying and versioning.
  • Own the agent runtime, including tool registries, permissions, function calling, grounding, retrieval, state, and long-running workflows.
  • Build training, scoring, experiment tracking, packaging, drift monitoring, retraining, and tuning components for classical ML and deep models.
  • Develop human-in-the-loop review and safe-actioning mechanisms before agents interact with dealer systems.
  • Evolve the domain graph and entity resolution while building reliable data ingestion and real-time context services with access controls and lineage.
  • Implement hybrid graph, vector, and keyword retrieval plus caching and TTL strategies.
  • Run offline and online evaluations for quality, factuality, bias, and safety.
  • Define latency, uptime, and cost SLOs while enabling autoscaling and spend controls.
  • Maintain model and agent registries, versioning, approvals, audit trails, reproducibility, and compliance support.
  • Provide templates, CLIs, sandboxes, and documentation, mentor engineers, and champion MLOps and AI safety practices.

Requirements

  • 5+ years building large-scale data, ML, or platform systems.
  • Strong software engineering fundamentals in API design, concurrency, and distributed systems.
  • Production experience with Python plus one of Java, Scala, or Go.
  • Experience with microservices, API design, MLOps pipelines, model tracking and registries, model CI/CD, A/B testing, shadowing, canarying, and online feature computation.
  • Experience with AWS, Docker, Kubernetes, and performance, reliability, and cost engineering for multi-tenant SaaS.
  • Practical ML knowledge covering feature engineering, training, evaluation, drift detection, and deployment of models supporting user-facing workflows.
  • Experience building or operating an LLM gateway or control plane with provider adapters, routing, policies, caching, quotas, rate limits, and cost and token accounting.
  • Experience with agentic systems, including tool use, function calling, orchestration, human-in-the-loop workflows, safety guardrails, and online evaluation and telemetry.
  • Experience with knowledge graphs, GraphQL, vector search, and hybrid retrieval patterns.
  • Ability to support developer experience, observability, access control, documentation, teaching, vendor portability, and cost-aware system design.

Tech Stack

Apache AirflowApache FlinkApache KafkaApache SparkAWSDockerGoGraphQLgRPCJavaKubernetesLightGBMMLflowNeo4jPythonScalaXGBoost

Categories

AI & MLBackendData Engineering