SambaNova Systems

Inference Systems Performance Architect

SambaNova Systems
Apply
4 hours ago
San Jose, CA, USAStaff+
H1B Sponsor

Base Salary

$245k - $325k/yr

Responsibilities

  • Define and drive the technical strategy for inference-systems performance across workload capture, benchmarking, modeling, and simulation.
  • Build workload-capture and agentic-benchmarking capabilities that represent production traffic and identify artificial contention or misleading cache-hit rates.
  • Own performance modeling and simulation practices to inform capacity planning, customer SLOs, and future systems and hardware planning.
  • Develop profiling tools that accurately localize bottlenecks across distributed inference hosts, accelerators, and fabrics.
  • Act as the senior technical voice across model optimization, systems, hardware, and product teams while evaluating reliability, scalability, operational cost, and adoption trade-offs.
  • Represent SambaNova’s performance capabilities to customers and partners.
  • Mentor principal and senior engineers and create systems, tools, and patterns that improve organizational productivity.
  • Drive resolution of ambiguous, novel, cross-organizational performance challenges.

Requirements

  • 12+ years of experience in performance engineering with a record of technical leadership on large-scale, complex systems.
  • Deep expertise in end-to-end performance analysis of distributed systems and bottleneck localization.
  • Proven expertise in realistic workload generation, simulation, and performance modeling calibrated against variable real-world workloads.
  • Ability to apply transferable performance methods in unfamiliar domains.
  • Ability to lead cross-functional efforts, mentor senior engineers, and influence organizational direction.
  • Experience credibly representing an organization to customers and partners.
  • Track record of independently scoping and delivering high-complexity, high-ambiguity work with significant product or roadmap impact.
  • Preferred: direct experience with LLM inference serving, continuous batching, prompt/KV caching, prefill/decode disaggregation, and tail-latency SLOs.
  • Preferred: familiarity with inference simulation frameworks or agentic benchmarking efforts.
  • Preferred: public technical presence through talks, writing, or community participation on systems performance.

Benefits

  • Base benefits include medical insurance with 95% employee premium coverage and 77% dependent premium coverage, plus an employer-contributed Health Savings Account.
  • Dental, vision, short- and long-term disability, basic life, voluntary life, AD&D, and flexible spending account options are available.
  • Well-being benefits include Headspace, Gympass+ with access to physical gyms, One Medical, counseling services, and an Employee Assistance Program.
  • The role is a full-time position for US-based employment.

Categories

SambaNova Systems

About SambaNova Systems

201-500 employees

Welcome to SambaNova: Revolutionizing AI Capacity At SambaNova, we're empowering developers, enterprises, governments, and data centers to unlock their full AI potential. Our full-stack infrastructure, from chips to models, enables lightning-fast performance, low power consumption, and high-efficiency computing. Our Mission To give every developer, enterprise, government and data center absolute sovereignty over their own data, models and AI infrastructure – to future-proof the AI workloads that will power and scale tomorrow. Our Technology We give our customers the optionality to experience SambaNova through the cloud or on-premise. Samba Cloud delivers the fastest inferences on the largest open source models like Llama 4 and DeepSeek. Developers can get started building in minutes with our OpenAI compatible APIs. All customers start on the developer tier and when they need more capacity can scale into our enterprise tier. SambaStack is our on-premise offering which includes the system, the platform, and foundation models. These components combine into a powerful technology stack that delivers unparalleled performance, ease of use, accuracy, data privacy, and the ability to power every use case across the world's largest organizations. SambaManaged is a modular and ready-to-deploy AI cloud designed to deliver unmatched efficiency for data centers and cloud service providers. This solution allows organizations to quickly deploy advanced AI inference services—without the need for costly infrastructure upgrades or specialized expertise—in as little as 90 days. At the heart of SambaNova innovation is the Reconfigurable Dataflow Unit (RDU). Purpose built for AI workloads, the RDU takes advantage of a dataflow architecture and a three-tiered memory design. The three tiers of memory enable the platform to run hundreds of models on a single node and to switch between them in microseconds. In 2023, SambaNova released its 4th generation RDU chip, the SN40L.