about 1 month ago
Responsibilities
- Develop a deep understanding of the AI Platform architecture, services, data and feature flows, and supported ML lifecycle.
- Own end-to-end quality for feature pipelines, training and serving infrastructure, model registries, deployment pipelines, LLM gateways, and vector or data stores.
- Design and implement automated testing frameworks for platform APIs, SDKs, services, and infrastructure components.
- Validate feature and model input/output correctness, completeness, and consistency throughout the platform.
- Test model-serving and inference services for correctness, latency, throughput, scalability, and reliability under load.
- Build regression and reconciliation suites for platform changes, migrations, and version upgrades.
- Validate ingestion, feature computation, model deployment and rollback, multi-tenancy, and governance capabilities.
- Define quality metrics and SLO validation with automated monitoring for platform anomalies.
- Contribute to evaluation frameworks for platform-level AI services, including serving correctness, feature integrity, and evaluation-infrastructure reliability.
- Use AI and LLMs to generate tests, synthetic data, and automation scaffolding.
- Drive root cause analysis for production platform issues and partner with engineering teams on preventive improvements.
- Champion quality engineering and reliability practices across the AI Platform organization.
Requirements
- 8+ years of experience in SDET, quality engineering, or software engineering, with strong experience building automation frameworks for platform, infrastructure, or backend services.
- Strong programming skills in Python and/or Java or Go for production-quality automation and tooling.
- Experience testing distributed systems, APIs, data pipelines, or ML/data platforms.
- Understanding of ML lifecycle components such as feature stores, training and serving, model registries, and deployment, or a strong platform-testing background with ML fluency.
- Experience with performance, load, and scalability testing.
- Familiarity with infrastructure-as-code, containers including Docker and Kubernetes, and cloud platforms including AWS, GCP, or Azure.
- Understanding of LLM and ML failure modes and evaluation concepts.
- Strong debugging and root cause analysis skills across service, data, and infrastructure layers.
- Experience with ML platform tools such as MLflow, Feast, KServe/Seldon, Ray, Kubeflow, SageMaker/Vertex AI, or LLM infrastructure such as vLLM, model gateways, and vector databases including Milvus or Pinecone.
- Experience with observability and monitoring tools such as Prometheus, Grafana, or Datadog.
- Experience testing multi-tenant, high-throughput SaaS platforms.
- Contributions to internal test frameworks or developer tooling.
- Exposure to chaos or reliability engineering.
- Excellent collaboration and communication skills.