
Inference Engineer
Hyperbolic Labs8 hours ago
Responsibilities
- Deploy and serve models across globally distributed clusters and heterogeneous hardware.
- Build inference capabilities on Forge and Kubernetes, including evaluating inference frameworks and serving engines.
- Set up production monitoring, gateways, and endpoints for inference services.
- Develop inference optimization, autoscaling, and KV-cache orchestration capabilities.
- Own inference delivery end to end, from initial implementation through serving real customer traffic.
- Debug customer inference and performance issues.
- Evaluate and apply deployment approaches across heterogeneous accelerators.
Requirements
- Broad inference expertise covering the full path from request to token.
- Deep hands-on Kubernetes experience operating production clusters.
- Understanding of TTFT, disaggregated inference, speculative decoding, and KV-cache internals.
- Familiarity with modern inference frameworks and serving engines, with the ability to evaluate and select among them.
- Working knowledge of NVIDIA Dynamo and distributed serving architectures.
- Experience setting up monitoring, gateways, and endpoints for production inference services.
- Demonstrated ability to build and launch a product or service serving real traffic.
- Experience spanning inference deployment and optimization is preferred.
- Hands-on model optimization experience, including quantization, batching strategies, or kernel-level tuning, is preferred.
- Understanding of RDMA and high-performance networking for distributed serving is preferred.
- Experience deploying inference across heterogeneous accelerators is preferred.
- Experience supporting customers with inference debugging and performance issues is preferred.
- Experience at a GPU cloud, inference provider, or AI infrastructure company is preferred.
Tech Stack
Categories
About Hyperbolic Labs
Hyperbolic Labs builds an open-access AI cloud that provides on-demand GPU clusters, a GPU marketplace for idle compute, and managed inference for AI startups, ML teams, and researchers. It sells capacity and services for training and serving models at production scale. The company is privately held, headquartered in San Francisco, and raised a Series A in 2024.