6 months ago
Cupertino, CA, USASenior
Responsibilities
- Contribute to the architecture and design of the Sohu host software stack
- Implement high-performance, modular code across the Rust, C++, and Python software stack
- Interface with firmware and driver teams to deliver a high-performance hardware/software stack
- Work with AI model researchers and product-facing teams on the Etched serving front-end
- Build scheduling logic for continuous batching and real-time inference
- Implement inference-time acceleration techniques including speculative decoding, tree search, and KV cache sharing
- Implement distributed networking primitives for efficient multi-server inference
Requirements
- Experience with C++ and Python
- Familiarity with transformer model architectures and inference serving stacks such as vLLM and SGLang, or experience in distributed inference or training environments
- Experience working cross-functionally in large software and hardware organizations
- Rust experience is a strong plus
- Familiarity with GPU kernels, the CUDA compilation stack, related tools, or other hardware accelerators is a strong plus
- Understanding of distributed systems, networking, and parallel programming is a strong plus
Benefits
- Full medical, dental, and vision coverage with 100% of premiums covered
- Housing subsidy for employees living within walking distance of the office
- Daily lunch and dinner provided at the office
- Relocation support for moves to Cupertino
- Fully in-person work in Cupertino
Categories
About Etched
Etched designs and sells AI inference chips and servers that hard‑code support for specific model architectures. Its first ASIC, Sohu, targets transformer models to deliver high‑throughput, low‑latency inference for workloads such as real‑time generation. Founded in 2022 and headquartered in San Jose, California, the privately held company serves organizations building large‑scale AI inference clusters and infrastructure.
