
Machine Learning Engineer — Inference Optimization
Featherless AI8 months ago
Remote, United StatesSenior
Responsibilities
- Optimize inference latency, throughput, and cost for large-scale ML models in production.
- Profile GPU and CPU inference pipelines to identify memory, kernel, batching, and I/O bottlenecks.
- Implement and tune quantization, KV-cache optimization and reuse, speculative decoding, batching, streaming, pruning, and architectural simplifications.
- Collaborate with research engineers to productionize new model architectures.
- Build and maintain inference-serving systems using Triton, custom runtimes, or bespoke stacks.
- Benchmark performance across NVIDIA and AMD GPUs, CPUs, and cloud setups.
- Improve reliability, observability, and cost efficiency under real workloads.
Requirements
- Strong experience in ML inference optimization or high-performance ML systems.
- Solid understanding of deep learning internals, including attention, memory layout, and compute graphs.
- Hands-on experience with PyTorch or a similar framework and model deployment.
- Familiarity with GPU performance tuning using CUDA, ROCm, Triton, or kernel-level optimizations.
- Experience scaling inference for real users rather than only research benchmarks.
- Experience with LLM or long-context model inference is desirable.
- Knowledge of TensorRT, ONNX Runtime, vLLM, or Triton is desirable.
- Experience optimizing across different hardware vendors is desirable.
- Open-source contributions in ML systems or inference tooling are desirable.
- Background in distributed systems or low-latency services is desirable.
Tech Stack
Categories
About Featherless AI
Featherless AI builds a serverless inference platform that orchestrates GPUs and load balances models so teams can deploy and scale open‑source AI without managing infrastructure. Its public cloud serves tens of thousands of open‑weight models and supports fine‑tuning, targeting developers, ML engineers, and enterprises needing reliable, high‑throughput inference. Founded in 2023 and headquartered in San Francisco, the privately held, Series A company is backed by investors including AMD and Airbus Ventures.