GrepJob
Tekion

Staff Software Development Test Engineer

Tekion
Apply
about 1 month ago
Bengaluru, IndiaStaff+
H1B Sponsor

Responsibilities

  • Develop a deep understanding of the AI Platform architecture, services, data and feature flows, and supported ML lifecycle.
  • Own end-to-end quality for feature pipelines, training and serving infrastructure, model registries, deployment pipelines, LLM gateways, and vector or data stores.
  • Design and implement automated testing frameworks for platform APIs, SDKs, services, and infrastructure components.
  • Validate feature and model input/output correctness, completeness, and consistency throughout the platform.
  • Test model-serving and inference services for correctness, latency, throughput, scalability, and reliability under load.
  • Build regression and reconciliation suites for platform changes, migrations, and version upgrades.
  • Validate ingestion, feature computation, model deployment and rollback, multi-tenancy, and governance capabilities.
  • Define quality metrics and SLO validation with automated monitoring for platform anomalies.
  • Contribute to evaluation frameworks for platform-level AI services, including serving correctness, feature integrity, and evaluation-infrastructure reliability.
  • Use AI and LLMs to generate tests, synthetic data, and automation scaffolding.
  • Drive root cause analysis for production platform issues and partner with engineering teams on preventive improvements.
  • Champion quality engineering and reliability practices across the AI Platform organization.

Requirements

  • 8+ years of experience in SDET, quality engineering, or software engineering, with strong experience building automation frameworks for platform, infrastructure, or backend services.
  • Strong programming skills in Python and/or Java or Go for production-quality automation and tooling.
  • Experience testing distributed systems, APIs, data pipelines, or ML/data platforms.
  • Understanding of ML lifecycle components such as feature stores, training and serving, model registries, and deployment, or a strong platform-testing background with ML fluency.
  • Experience with performance, load, and scalability testing.
  • Familiarity with infrastructure-as-code, containers including Docker and Kubernetes, and cloud platforms including AWS, GCP, or Azure.
  • Understanding of LLM and ML failure modes and evaluation concepts.
  • Strong debugging and root cause analysis skills across service, data, and infrastructure layers.
  • Experience with ML platform tools such as MLflow, Feast, KServe/Seldon, Ray, Kubeflow, SageMaker/Vertex AI, or LLM infrastructure such as vLLM, model gateways, and vector databases including Milvus or Pinecone.
  • Experience with observability and monitoring tools such as Prometheus, Grafana, or Datadog.
  • Experience testing multi-tenant, high-throughput SaaS platforms.
  • Contributions to internal test frameworks or developer tooling.
  • Exposure to chaos or reliability engineering.
  • Excellent collaboration and communication skills.