
Senior Software Development Engineer in Test (SDET) - AI Cluster Networking and Security
Cerebras Systems2 months ago
Bengaluru, IndiaSenior
Responsibilities
- Define optimized test strategies and methodologies for cutting-edge AI infrastructure.
- Build and maintain automated tests for cluster features covering high availability, failure scenarios, performance, stress, and security.
- Break down large distributed ML training and inference systems into components suitable for unit testing.
- Test AI cluster software, hardware components, high-speed interconnects, and data-transfer systems.
- Qualify networking solutions, including high-speed switches, routers, optics, and vendor platforms.
- Validate cluster security features including operating system security, network security, cloud compliance, user access, and security certifications.
- Champion cluster reliability, security, observability, and uptime targets up to 99.9999%.
Requirements
- Bachelor’s or master’s degree in engineering, computer science, electrical engineering, AI, data science, or a related field.
- 10+ years of experience testing enterprise software, distributed systems, datacenter hardware, or datacenter software.
- Experience with enterprise or cloud networking infrastructure, high-speed switches, routers, and firewalls.
- Experience qualifying Juniper, Arista, or Cisco networking platforms and Ixia or Spirent network test equipment.
- Experience with datacenter technologies and protocols including BGP, ECN, and PFC.
- Experience testing networking security, compliance, and firewalls.
- Strong coding skills in Python, Golang, or C/C++.
- Strong debugging skills for large distributed systems, hardware, and software, including experience with tools such as gdb, strace, and networking monitors.
- Strong understanding of operating system internals, memory management, file systems, security basics, performance, datacenter layout, PCIe, networking, and storage.
- Experience with AWS, Kubernetes, and Docker; Grafana and Prometheus experience is a strong plus.
- Understanding of ML model training and inference is a strong plus.
- Understanding and experience with ML hardware accelerators such as GPUs and custom accelerator ASICs is a strong plus.
Benefits
- Opportunity to build a breakthrough AI platform beyond GPU constraints.
- Opportunity to publish and open source cutting-edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Job stability with startup vitality.
- Simple, non-corporate work culture that respects individual beliefs.
- Equal and diverse work environment with continuous learning, growth, and support.
About Cerebras Systems
Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.