
AI Inference Core - Infrastructure SW Engineer
Cerebras Systems1 day ago
Sunnyvale, CA, USAMid Level
H1B sponsor
Responsibilities
- Design, develop, test, and maintain Python frameworks and services that orchestrate engineering workflows across machines and clusters.
- Build reusable abstractions for scheduling, distributed execution, resource management, test execution, workflow planning, and failure recovery.
- Define maintainable APIs, module boundaries, extension points, and data models for evolving infrastructure systems.
- Handle concurrency, asynchronous execution, multiprocessing, state management, retries, idempotency, cancellation, and partial failures.
- Debug complex issues across Python applications, operating systems, processes, filesystems, networking, remote machines, and distributed services.
- Write automated tests and documentation for reliable, reusable infrastructure.
- Partner with platform, CI, release, quality, ML systems, and product engineering teams to translate requirements into scalable software designs.
Requirements
- At least 3 years of professional software-engineering experience.
- Strong proficiency in Python and understanding of its strengths, limitations, and runtime behavior.
- Experience designing maintainable software systems, libraries, frameworks, backend services, or developer-facing APIs.
- Knowledge of software architecture, abstraction boundaries, design patterns, extensibility, and long-term maintainability.
- Understanding of processes, threads, asynchronous execution, synchronization, and shared state.
- Foundational understanding of operating-system concepts including processes, signals, filesystems, resource management, and program execution.
- Foundational understanding of distributed-systems concepts including retries, timeouts, idempotency, partial failure, coordination, and eventual consistency.
- Strong debugging and independent problem-solving skills.
- Preferred experience with Python concurrency technologies such as asyncio, multiprocessing, concurrent futures, or event-driven systems.
- Preferred experience building orchestration engines, workflow systems, schedulers, distributed job runners, control-plane software, or test infrastructure.
- Preferred familiarity with CI systems, build systems, release infrastructure, developer-productivity tooling, Kubernetes, containerized environments, cluster schedulers, or remote execution systems.
- BS or MS in Computer Science or a related field, or equivalent practical experience.
Benefits
- Opportunity to build an AI platform beyond GPU constraints and work on a fast AI supercomputer.
- Opportunities to publish and open source AI research.
- Job stability with startup vitality and a non-corporate work culture.
- Cerebras is committed to an equal and diverse environment and supports continuous learning, growth, and team support.
Tech Stack
Categories
About Cerebras Systems
Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.