SambaNova Systems

Principal AI Systems Performance Engineer

SambaNova Systems
Apply
3 hours ago
San Jose, CA, USAStaff+
H1B Sponsor

Responsibilities

  • Bring up and optimize foundation models including DeepSeek, Llama, and Qwen on the SambaNova platform and software stack.
  • Profile and improve model performance across compiler, runtime, and hardware layers to achieve state-of-the-art throughput and latency.
  • Collaborate with machine learning, compiler, runtime, and hardware teams on co-designed, high-performance AI applications.
  • Integrate advances in model architecture, quantization, scheduling, and memory optimization.
  • Develop scalable, efficient end-to-end inference solutions aligned with customer needs.
  • Identify performance bottlenecks and propose dataflow or scheduling optimizations for single-node and distributed systems.

Requirements

  • Bachelor’s or higher degree in computer science, electrical engineering, or a related field such as applied mathematics, physics, or statistics.
  • At least 3 years of experience in deep-learning model development and performance optimization, compiler/runtime/kernel optimization, software-hardware co-design, or systems performance tuning.
  • Proficiency in Python or C++ with strong foundations in algorithms, data structures, and numerical computing.
  • Experience with at least one major ML framework: PyTorch, TensorFlow, or JAX.
  • Demonstrated ability to analyze and optimize performance in real-world ML pipelines.
  • Preferred experience includes LLM or multimodal model training and inference, distributed training, continuous batching, high-throughput inference, quantization, graph optimization, kernel fusion, and model partitioning.
  • Preferred familiarity with DeepSpeed, Megatron, vLLM, TensorRT, CUDA, Triton, OpenCL, cuDNN, or cuBLAS.
  • Knowledge of memory hierarchy optimization, caching, and scheduling for large-scale model execution is preferred.
  • Publication records or open-source contributions in ML systems or performance optimization are a plus.

Benefits

  • US-based full-time employees receive medical insurance with 95% employee premium coverage and 77% dependent premium coverage.
  • Health Savings Account with employer contribution, Dental, Vision, short- and long-term disability, Basic Life, Voluntary Life, AD&D, and Flexible Spending Account options are offered.
  • Well-being benefits include Headspace, Gympass+ with physical gym access, One Medical, counseling, and an Employee Assistance Program.

Tech Stack

Categories

SambaNova Systems

About SambaNova Systems

201-500 employees

Welcome to SambaNova: Revolutionizing AI Capacity At SambaNova, we're empowering developers, enterprises, governments, and data centers to unlock their full AI potential. Our full-stack infrastructure, from chips to models, enables lightning-fast performance, low power consumption, and high-efficiency computing. Our Mission To give every developer, enterprise, government and data center absolute sovereignty over their own data, models and AI infrastructure – to future-proof the AI workloads that will power and scale tomorrow. Our Technology We give our customers the optionality to experience SambaNova through the cloud or on-premise. Samba Cloud delivers the fastest inferences on the largest open source models like Llama 4 and DeepSeek. Developers can get started building in minutes with our OpenAI compatible APIs. All customers start on the developer tier and when they need more capacity can scale into our enterprise tier. SambaStack is our on-premise offering which includes the system, the platform, and foundation models. These components combine into a powerful technology stack that delivers unparalleled performance, ease of use, accuracy, data privacy, and the ability to power every use case across the world's largest organizations. SambaManaged is a modular and ready-to-deploy AI cloud designed to deliver unmatched efficiency for data centers and cloud service providers. This solution allows organizations to quickly deploy advanced AI inference services—without the need for costly infrastructure upgrades or specialized expertise—in as little as 90 days. At the heart of SambaNova innovation is the Reconfigurable Dataflow Unit (RDU). Purpose built for AI workloads, the RDU takes advantage of a dataflow architecture and a three-tiered memory design. The three tiers of memory enable the platform to run hundreds of models on a single node and to switch between them in microseconds. In 2023, SambaNova released its 4th generation RDU chip, the SN40L.