3 months ago
Bengaluru, IndiaStaff+
Responsibilities
- Develop a deep understanding of the AI Platform architecture, services, data and feature flows, and supported ML lifecycle.
- Own end-to-end quality for feature pipelines, training and serving infrastructure, model registries, deployment pipelines, LLM gateways, and vector or data stores.
- Design and implement automated testing frameworks for platform APIs, SDKs, services, and infrastructure components.
- Validate feature and model input/output correctness, completeness, and consistency across platform flows.
- Test model-serving and inference services for correctness, latency, throughput, scalability, and reliability under load.
- Build regression and reconciliation suites for platform changes, migrations, and version upgrades.
- Validate ingestion, feature computation, model deployment and rollback, multi-tenancy, and governance capabilities.
- Define quality metrics and SLO validation with automated monitoring for platform anomalies.
- Contribute to evaluation frameworks for platform-level AI services and use AI or LLMs to accelerate test generation, synthetic data creation, and automation scaffolding.
- Drive root cause analysis for production platform issues and champion quality and reliability improvements across the AI Platform organization.
Requirements
- 8+ years of experience in SDET, quality engineering, or software engineering, including automation frameworks for platform, infrastructure, or backend services.
- Strong programming skills in Python and/or Java or Go for production-quality automation and tooling.
- Experience testing distributed systems, APIs, data pipelines, or ML and data platforms.
- Understanding of ML lifecycle components such as feature stores, training and serving, model registries, and deployment.
- Experience with performance, load, and scalability testing.
- Familiarity with containers including Docker and Kubernetes and cloud platforms including AWS, GCP, or Azure.
- Understanding of LLM and ML failure modes and evaluation concepts.
- Strong debugging and root cause analysis skills across service, data, and infrastructure layers.
- Experience with ML platform tools such as MLflow, Feast, KServe, Seldon, Ray, Kubeflow, SageMaker, or Vertex AI is preferred.
- Experience with LLM infrastructure such as vLLM, model gateways, or vector databases such as Milvus and Pinecone is preferred.
- Experience with observability and monitoring tools such as Prometheus, Grafana, or Datadog is preferred.
- Experience testing multi-tenant, high-throughput SaaS platforms, contributing to internal test frameworks or developer tooling, or exposure to chaos or reliability engineering is preferred.
Tech Stack
Categories
About Tekion
Tekion builds an AI-native, cloud platform for automotive retail that unifies dealers, OEMs, and partners. Its products include Automotive Retail Cloud (a dealership management system for retailers), Automotive Enterprise Cloud for manufacturers, and Automotive Partner Cloud for integrations, delivered as subscription software. Privately held and headquartered in Pleasanton, California, Tekion raised private equity funding in 2024.
