GrepJob
Tekion

Staff Machine Learning Engineer

Tekion
Apply
7 months ago
Bengaluru, IndiaStaff+
H1B Sponsor

Responsibilities

  • Convert tabular, NLP/LLM, recommendation, and forecasting models into fast, reliable production services with defined API contracts.
  • Integrate with the LLM Gateway/MCP and manage prompt and configuration versioning.
  • Build and orchestrate CI/CD pipelines and develop reusable templates, SDKs, CLIs, sandbox datasets, and documentation.
  • Review, refactor, optimize, containerize, deploy, version, and monitor data science models.
  • Design batch and stream pipelines, online features, and domain-graph integrations.
  • Build inference microservices with REST/gRPC, schema versioning, structured outputs, and stringent p95 latency targets.
  • Manage feature stores, model and agent registries, versioning, lineage, and the overall model/feature lifecycle.
  • Implement prompt versioning, A/B testing, guardrails, and feedback- and metrics-based orchestration.
  • Instrument observability for traces, logs, metrics, drift, model performance, safety signals, and cost tracking.
  • Ensure runtime reliability through autoscaling, caching, circuit breakers, retries, fallbacks, and graceful degradation.
  • Collaborate with Applied Sciences, data engineering, product management, and architecture teams to deliver enterprise ML systems.
  • Monitor and mitigate risks associated with LLMs and agentic systems.

Requirements

  • 8–11+ years of experience in ML engineering, MLOps, or backend/platform engineering with production ML.
  • Experience with LLMs, retrieval systems, vector stores, and graph or knowledge stores.
  • Strong Python skills plus one of Java, Go, or Scala, along with API design, concurrency, and testing experience.
  • Hands-on experience with orchestration frameworks and libraries such as LangChain, LlamaIndex, OpenAI Function Calling, or AgentKit.
  • Knowledge of reactive, planning, and retrieval-augmented agent architectures and safe execution patterns.
  • Experience with Airflow or Kubeflow, Spark or Flink, Kafka or Kinesis, and strong data quality practices.
  • Experience with Docker, Kubernetes, service meshes, REST/gRPC, and performance and reliability engineering.
  • Experience with experiment tracking, model registries such as MLflow, feature stores, A/B testing, shadow testing, and drift detection.
  • Experience with OpenTelemetry, Prometheus, and Grafana, including diagnosing latency, tail behavior, and memory or CPU hotspots.
  • Experience with AWS services including IAM, ECS/EKS, S3, RDS/DynamoDB, and Step Functions/Lambda, with cost optimization experience.
  • Knowledge of secrets management, RBAC/ABAC, PII handling, and auditability.
  • Preferred: product-oriented, reliability- and safety-first, systems-thinking, collaborative, and pragmatic approach.

Benefits

  • Opportunity to work on a global, cloud-native, AI-driven automotive platform.
  • Collaborative, innovative, and fast-paced work environment.
  • Current Tekion employees should apply through the internal Ashby job board effective August 4, 2026.

Tech Stack

Amazon DynamoDBApache AirflowApache FlinkApache KafkaApache SparkAWSDockerGoGrafanagRPCJavaKubernetesMLflowPrometheusPythonScala

Categories

AI & MLBackendData Engineering