Staff Software Engineer - Tools & Infrastructure / DevOps
Cerebras SystemsResponsibilities
- Design, build, and evolve CI/CD pipelines for reliable build, test, and release workflows.
- Own artifact lifecycle systems covering versioning, storage, distribution, dependency management, and reproducible builds.
- Improve code review workflows, branching strategies, repository management, and automated integration processes.
- Provision, monitor, and optimize AWS or other cloud infrastructure supporting CI workloads.
- Troubleshoot build failures, pipeline bottlenecks, and infrastructure issues through root-cause analysis and durable fixes.
- Improve internal build infrastructure, test infrastructure, developer tooling, and automation to increase engineering productivity.
- Lead architectural improvements to developer infrastructure and identify systemic workflow bottlenecks.
- Contribute to AI tooling that improves engineering productivity and automates repetitive workflows.
- Provide technical leadership, influence engineering standards, and mentor engineers.
- Participate in on-call and incident response for Developer Productivity systems and services.
Requirements
- 7+ years of professional experience in software engineering, infrastructure engineering, DevOps, developer productivity, or a related area.
- Deep hands-on experience with CI/CD systems and automated build, test, and deployment infrastructure.
- Experience with artifact repositories, software packaging, dependency management, and reproducible builds.
- Experience with cloud computing platforms and programmatic infrastructure provisioning; AWS is preferred.
- Strong experience with distributed version control, code review workflows, branching strategies, and repository management.
- Strong understanding of Linux/Unix systems, networking fundamentals, and scripting or programming for automation.
- Experience with containerization and container orchestration; Kubernetes is preferred.
- Experience troubleshooting complex distributed systems and infrastructure issues.
- Demonstrated experience leading technical initiatives across multiple teams and improving developer infrastructure at organizational scale.
- Ability to identify architectural bottlenecks, evaluate tradeoffs, and drive long-term improvements.
- Experience operating production systems and participating in on-call, incident response, and postmortem processes.
- Experience with infrastructure-as-code tools and practices is preferred.
- Proficiency in Python, Go, Shell, or another language used for infrastructure automation and developer tooling is preferred.
- Experience with build systems, build graph optimization, or large-scale build infrastructure is preferred.
- Experience with observability, including monitoring, logging, alerting, and performance analysis, is preferred.
- Experience building internal developer platforms, self-service tooling, or developer-facing infrastructure is preferred.
- Experience applying AI/LLM tooling to engineering workflows or developer productivity is preferred.
- BS/MS in Computer Science or a related field, or equivalent practical experience.
Benefits
- Opportunity to build infrastructure for a breakthrough AI platform and work with a high-performance AI supercomputer.
- Opportunities to publish and open source AI research.
- Job stability with startup vitality and a simple, non-corporate work culture.
- Inclusive, equal-opportunity workplace emphasizing continuous learning, growth, and support.
About Cerebras Systems
Cerebras Systems is the world's fastest AI inference. We are powering the future of generative AI. We’re a team of pioneering computer architects, deep learning researchers, and engineers building a new class of AI supercomputers from the ground up. Our flagship system, Cerebras CS-3, is powered by the Wafer Scale Engine 3—the world’s largest and fastest AI processor. CS-3s are effortlessly clustered to create the largest AI supercomputers on Earth, while abstracting away the complexity of traditional distributed computing. From sub-second inference speeds to breakthrough training performance, Cerebras makes it easier to build and deploy state-of-the-art AI—from proprietary enterprise models to open-source projects downloaded millions of times. Here’s what makes our platform different: 🔦 Sub-second reasoning – Instant intelligence and real-time responsiveness, even at massive scale ⚡ Blazing-fast inference – Up to 100x performance gains over traditional AI infrastructure 🧠 Agentic AI in action – Models that can plan, act, and adapt autonomously 🌍 Scalable infrastructure – Built to move from prototype to global deployment without friction Cerebras solutions are available in the Cerebras Cloud or on-prem, serving leading enterprises, research labs, and government agencies worldwide. 👉 Learn more: www.cerebras.ai Join us: https://cerebras.net/careers/