
Senior SDET, Inference Platform
Cerebras Systems2 months ago
Responsibilities
- Design, build, and maintain test infrastructure and automation for deploying and validating the Cerebras Inference Platform.
- Validate the platform across cloud-managed Kubernetes environments and deployments on Cerebras hardware.
- Test Kubernetes workloads, CI/CD pipelines, ingress, service discovery, NGINX, and load balancing.
- Debug complex issues across networking, orchestration, deployment, and distributed services.
- Develop testbeds for platform performance, scalability, reliability, and inference performance.
- Contribute to test plans and validation strategies for new platform features and releases.
- Improve observability, diagnostics, and debugging workflows across the platform stack.
- Collaborate with platform and engineering teams to deliver reliable, production-ready releases.
- Mentor junior engineers.
Requirements
- At least 3 years of experience in software engineering, QA or quality engineering, systems engineering, or infrastructure development.
- Strong programming skills in Python and/or Go; experience with both is a plus.
- Experience building automation tools, testing frameworks, or internal developer tooling.
- Hands-on experience with CI/CD systems such as Jenkins.
- Experience debugging complex systems, distributed services, or networked infrastructure.
- Familiarity with systems-level development, infrastructure tooling, or platform integration.
- Strong problem-solving, communication, and collaboration skills.
- Hands-on experience with Kubernetes and container orchestration in production or staging environments is preferred.
- Experience with Amazon EKS, ArgoCD, ingress controllers, service discovery, NGINX, load balancing, Bazel, or k9s is preferred.
- Exposure to performance debugging, profiling, system observability tools, ML inference infrastructure, model serving systems, or GPU-accelerated workloads is preferred.
Benefits
- Role based in Toronto or Sunnyvale.
- Opportunity to work on a large-scale AI inference platform and Cerebras hardware.
- Opportunities to publish and open source AI research.
- Equal and diverse work environment with continuous learning, growth, and support.
Tech Stack
About Cerebras Systems
Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.