
Inference Engineer
Designworks Talent2 months ago
Bellevue, WA, USASenior / Staff+
Responsibilities
- Build and operate production-grade model-serving and inference systems for high-throughput, low-latency AI workloads.
- Optimize inference infrastructure for token throughput, latency, scalability, and cost efficiency across model architectures and workloads.
- Design systems that maximize GPU utilization while maintaining predictable performance and reliability.
- Improve the scalability and operational maturity of the inference platform as customer demand grows.
- Partner with AI training, GPU performance, orchestration, and infrastructure teams to transition models from development to production serving.
- Develop monitoring, alerting, and operational practices for reliable inference services.
- Investigate and resolve performance, reliability, and capacity challenges across inference workloads.
- Contribute to architecture decisions and engineering standards as the platform evolves.
Requirements
- Experience building and operating production machine-learning inference or model-serving systems at scale.
- Strong understanding of latency, throughput, memory utilization, and cost-efficiency trade-offs when serving large AI models.
- Experience designing reliable distributed systems or production infrastructure.
- Understanding of GPU-backed AI workloads and the challenges of scaling inference systems.
- Strong engineering fundamentals and the ability to independently own complex technical problems.
- Comfort working in a fast-moving environment where systems and processes are being built from the ground up.
- Preferred experience with inference-serving frameworks such as vLLM, TensorRT-LLM, or Triton Inference Server.
- Preferred experience optimizing LLM inference workloads or large-scale AI serving platforms.
- Preferred background operating API-based AI products or high-volume production services.
- Preferred experience with GPU scheduling, distributed systems, Kubernetes, or cloud infrastructure platforms.
- Familiarity with model optimization techniques including quantization, batching, caching, or performance tuning.
- Experience at a hyperscaler, AI lab, GPU cloud provider, or large-scale machine-learning infrastructure organization is preferred.
- U.S. work authorization is required; visa sponsorship is not currently available.
Benefits
- Hybrid role in the Bellevue, Washington area with approximately three days per week in the office.
- Candidates elsewhere in the U.S. may relocate; U.S. work authorization is required and visa sponsorship is unavailable.
- U.S.-based employees receive medical, dental, and vision insurance, a 401(k) plan with company match, and paid holidays.
- Eligible roles may receive merit increases, annual bonuses, and stock awards.
Tech Stack
Categories
About Designworks Talent
Designworks Talent specializes in the design and delivery of enterprise workforce solutions including strategy, talent acquisition, engagement, and succession planning. We work collaboratively with business leaders and enterprise Talent Acquisition and Human Resources teams to deliver custom workforce solutions to meet the dynamic talent demands of your business.