
Software Engineer, ML Serving - Rime Ai
Unusual Ventures2 months ago
Responsibilities
- Architect and implement Rime’s TTS serving infrastructure from GPU-backed inference engines through the API layer.
- Optimize model serving from single-node deployments to disaggregated fleets.
- Support NVIDIA hardware architectures from Hopper through Blackwell for on-premises and cloud deployments.
- Own continuous integration and deployment workflows for the model-serving pipeline.
- Participate in on-call rotations and maintain monitoring, alerting, and observability across the serving stack.
- Provision resources and manage costs across the GPU fleet.
Requirements
- Hands-on experience with real-time multinode ML serving infrastructure and frameworks such as NVIDIA Dynamo, Triton, vLLM, SGLang, or equivalent.
- Experience with distributed or disaggregated model serving, including Tensor Parallel or Pipeline Parallel approaches.
- Strong cloud infrastructure fundamentals, including Linux internals, networking, Docker, and Kubernetes.
- Experience with infrastructure as code using Terraform, Packer, or comparable tooling.
- Willingness to participate in on-call and share responsibility for production reliability.
- Preferred experience with multinode training using DDP or FSDP.
- Preferred experience with gRPC, bidirectional binary streaming, audio streaming, WebRTC, WebSockets, multilingual monorepos, multi-cloud infrastructure, and configuration management tools such as Ansible, Chef, or Puppet.
- SRE, DevOps, or platform engineering experience at a startup and experience at an early-stage company are preferred.
Benefits
- Meaningful equity upside at an early-stage company.
- Work location is San Francisco / Bay Area.
- Direct collaboration with inference, platform, and ML teams in a high-ownership, low-bureaucracy environment.