
Inference Engineer
Designworks Talentabout 5 hours ago
Bellevue, WA, USASenior / Staff+
Responsibilities
- Build and operate production-grade model-serving and inference systems.
- Optimize inference infrastructure for token throughput, latency, scalability, and cost efficiency.
- Design systems to maximize GPU utilization while ensuring performance and reliability.
- Improve scalability and operational maturity of inference platforms.
- Collaborate with AI training, GPU performance, and infrastructure teams.
- Develop monitoring and operational practices for reliable inference services.
- Investigate and resolve performance and capacity challenges.
- Contribute to architecture decisions and engineering standards.
Requirements
- Experience building and operating production machine learning inference systems at scale.
- Strong understanding of performance trade-offs in serving large AI models.
- Experience designing reliable distributed systems or production infrastructure.
- Understanding of GPU-backed AI workloads and scaling challenges.
- Strong engineering fundamentals and ability to own complex technical problems.
- Comfortable in a fast-moving environment with evolving systems.
Benefits
- Competitive base pay for the Bellevue market.
- Eligibility for merit increases, annual bonuses, and stock based on performance.
- Access to medical, dental, and vision insurance.
- 401(k) plan with company match.
- Paid holidays each calendar year.