Tekion

Staff Machine Learning Engineer

Tekion
Apply
9 months ago
Bengaluru, IndiaStaff+

Responsibilities

  • Convert tabular, NLP/LLM, recommendation, and forecasting models into fast, reliable production services with defined API contracts.
  • Integrate with the LLM Gateway/MCP and manage prompt and configuration versioning, A/B testing, guardrails, and dynamic orchestration.
  • Build and orchestrate CI/CD pipelines and develop templates, SDKs, CLIs, sandbox datasets, and documentation for ML delivery.
  • Review, refactor, optimize, containerize, deploy, version, and monitor data science models for quality.
  • Design batch and streaming pipelines, online features, and domain-graph integrations.
  • Build inference microservices with REST/gRPC, schema versioning, structured outputs, and stringent p95 latency targets.
  • Manage feature stores, model and agent registries, experiment lineage, versioning, and the broader model and feature lifecycle.
  • Implement observability for traces, logs, metrics, feature and data drift, model performance, safety signals, and cost tracking.
  • Ensure runtime reliability through autoscaling, caching, circuit breakers, retries, fallbacks, and graceful degradation.
  • Collaborate with Applied Sciences, data engineering, product management, and architecture teams to deliver measurable dealer and consumer outcomes.
  • Monitor and mitigate risks associated with LLMs and agentic systems while operationalizing secure, compliant, and explainable services.

Requirements

  • 8–11+ years of experience in ML engineering, MLOps, or backend/platform engineering with production ML.
  • Experience with LLMs, retrieval systems, vector stores, and graph or knowledge stores.
  • Strong software engineering fundamentals with Python and one of Java, Go, or Scala, plus API design, concurrency, and testing experience.
  • Hands-on experience with orchestration frameworks and libraries such as LangChain, LlamaIndex, OpenAI Function Calling, or AgentKit.
  • Knowledge of reactive, planning, and retrieval-augmented agent architectures and safe execution patterns.
  • Experience with Airflow or Kubeflow, Spark or Flink, Kafka or Kinesis, and strong data quality practices.
  • Experience with Docker, Kubernetes, service meshes, REST/gRPC, performance engineering, and reliability engineering.
  • Experience with experiment tracking, model registries such as MLflow, feature stores, A/B testing, shadow testing, and drift detection.
  • Experience with OpenTelemetry, Prometheus, and Grafana, including debugging latency, tail behavior, and memory or CPU hotspots.
  • Experience with AWS infrastructure and services, including IAM, ECS/EKS, S3, RDS/DynamoDB, Step Functions, and Lambda, plus cost optimization.
  • Knowledge of secrets management, RBAC/ABAC, PII handling, security, compliance, and auditability.
  • Preferred qualifications include product orientation, reliability and safety focus, systems thinking, collaboration, documentation, teaching, and pragmatic automation.

Benefits

  • Opportunity to work on complex challenges at scale across a global cloud-native, AI-driven platform.
  • Innovative, collaborative, and fast-paced work culture.

Tech Stack

Amazon DynamoDBApache AirflowApache FlinkApache KafkaApache SparkAWSDockerGoGrafanagRPCJavaKubernetesMLflowPrometheusPythonScala
Tekion

About Tekion

1,001-5,000 employees

Tekion builds an AI-native, cloud platform for automotive retail that unifies dealers, OEMs, and partners. Its products include Automotive Retail Cloud (a dealership management system for retailers), Automotive Enterprise Cloud for manufacturers, and Automotive Partner Cloud for integrations, delivered as subscription software. Privately held and headquartered in Pleasanton, California, Tekion raised private equity funding in 2024.

Contact me