Cerebras Systems

Sr. Member of Technical Staff

Cerebras Systems
Apply
4 months ago
Sunnyvale, CA, USASenior
H1B sponsor

Responsibilities

  • Design and develop fault-tolerant software features for resiliency and high availability across distributed environments.
  • Build and maintain AWS-based deployment workflows for low-latency, scalable AI inference services.
  • Develop Python scripts and APIs for data preprocessing, inference execution, and post-processing.
  • Use parallel programming techniques such as multithreading and asynchronous processing on AWS compute instances.
  • Develop performance visualization and analysis components for inference services.
  • Containerize inference software with Docker and define Kubernetes orchestration strategies for reliable scaling.
  • Create automated scripts to detect and mitigate common software failure modes.
  • Debug model deployment, container orchestration, and networking issues and document root causes.
  • Analyze logs, metrics, and distributed traces to triage and resolve service defects.
  • Collaborate with Product Management and User Experience teams on inference-service interfaces and requirements.
  • Write technical documentation for infrastructure configurations, inference workflows, and APIs.
  • Track defects, enhancements, and release notes using Jira and Git.

Requirements

  • Master's degree or foreign equivalent in Computer Science or a related field.
  • At least 18 months of experience as an Information Security Analyst, Software Engineer, Sr. Member of Technical Staff, IT Senior Applications Engineer, or in a related occupation.
  • Experience with infrastructure-as-code and deployment automation using Terraform, AWS CloudFormation, AWS CDK, and Ansible.
  • Experience with Docker, Kubernetes, AWS EKS, AWS Elastic Container Service (ECS), AWS Fargate, and Helm.
  • Experience with AWS EC2, AWS Lambda, and Auto Scaling Groups.
  • Experience with AWS CloudWatch, AWS X-Ray, ELK, Elasticsearch, Logstash, Kibana, Prometheus, and Grafana for monitoring, logging, and tracing.
  • Programming experience with Python, Node.js, JavaScript, and Flask.
  • Experience with PostgreSQL, Redis, and NFS.
  • Experience with Jenkins and Git for CI/CD and version control.

Benefits

  • Opportunity to build a breakthrough AI platform beyond GPU constraints.
  • Opportunity to publish and open source cutting-edge AI research.
  • Work on a high-performance AI supercomputer.
  • Job stability with startup vitality.
  • A simple, non-corporate culture that respects individual beliefs.
  • Inclusive work environment focused on continuous learning, growth, and support.

Tech Stack

AnsibleDockerElasticsearchFlaskGitGrafanaHelmJavaScriptJenkinsKibanaKubernetesLogstashNode.jsPostgreSQLPrometheusPythonRedisTerraform

Categories

Cerebras Systems

About Cerebras Systems

1,001-5,000 employees

Cerebras Systems designs and sells AI compute systems built around its wafer-scale WSE-3 processor, delivered as the CS-3 appliance and via the Cerebras Cloud. It targets enterprises, model labs, and government users needing fast training and inference, and offers on‑prem and cloud deployments. Privately held and headquartered in Sunnyvale, California, the company announced a multi-year partnership with OpenAI to deploy large-scale inference capacity.

Contact me