3 months ago
San Jose, CA, USASenior
Base Salary
$192k - $250k/yr
Responsibilities
- Design and implement high-performance, scalable inference systems for LLMs and VLMs across cloud, edge, and hybrid edge-cloud platforms.
- Develop and optimize custom kernels and operators for GPU, NPU, DSP, and other hardware accelerators.
- Integrate KV-cache management, tensor and model parallelism, quantization, and memory-efficient execution into production inference systems.
- Partner with systems and hardware teams to optimize hardware-software integration across diverse compute environments.
- Translate architectural requirements into robust, maintainable software meeting performance, safety, and reliability standards.
- Define and drive the roadmap for LLM/VLM inference in the AIOS stack.
- Monitor industry and competitor developments and apply relevant AI and large-scale systems engineering practices.
Requirements
- At least 5 years of hands-on software development experience building and optimizing AI inference systems at scale.
- Direct experience with LLM/VLM model internals, Transformer architectures, inference bottlenecks, and optimization techniques.
- Strong expertise in kernel development, parallelism, memory optimization, and distributed inference systems.
- Proficiency with GPU/NPU programming, CUDA or vendor-specific SDKs, compiler toolchains, and PyTorch or TensorFlow.
- Strong C/C++ programming skills and experience delivering high-performance production software.
- Strong foundation in computer architecture, systems programming, CPU/GPU pipelines, memory hierarchy, scheduling, and embedded systems.
- Bachelor's or master's degree in Computer Science, Computer Engineering, or a related technical field.
- Master's or PhD degree and 5 years of industry experience are preferred.
- Experience with inference serving systems for large models, including batching, scheduling, caching, and load balancing, is preferred.
- Expertise in hardware-aware optimization such as kernel fusion, mixed precision, quantization, and pruning is preferred.
- Familiarity with edge and embedded AI, real-time constraints, and limited-resource optimization is preferred.
- Contributions to widely used AI frameworks, libraries, or performance-critical software are preferred.
- Strong communication and cross-functional collaboration skills.
Benefits
- Full-time employees are eligible for medical plans including Anthem Blue Cross, HSA, and Kaiser HMO, with $0 employee-only coverage.
- Dental and vision plans offer options with $0 paycheck contribution for employees and eligible dependents.
- Company-paid HSA contributions are available with the High Deductible Anthem Blue Cross plan.
- Healthcare and dependent care FSAs, 401(k) with BrokerageLink, company-paid life and disability insurance, and an Employee Assistance Program are provided.
- Benefits include sick and vacation time, 13 paid holidays, paid parental leave, and paid disability leave subject to eligibility periods.
- Additional benefits include voluntary life and AD&D insurance, pet insurance, commuter benefits, mobile phone credit, free lunch and snacks, an onsite gym, and employee discounts and perks.
