GrepJob
Designworks Talent

Inference Engineer

Designworks Talent
Apply
about 5 hours ago
Bellevue, WA, USASenior / Staff+

Responsibilities

  • Build and operate production-grade model-serving and inference systems.
  • Optimize inference infrastructure for token throughput, latency, scalability, and cost efficiency.
  • Design systems to maximize GPU utilization while ensuring performance and reliability.
  • Improve scalability and operational maturity of inference platforms.
  • Collaborate with AI training, GPU performance, and infrastructure teams.
  • Develop monitoring and operational practices for reliable inference services.
  • Investigate and resolve performance and capacity challenges.
  • Contribute to architecture decisions and engineering standards.

Requirements

  • Experience building and operating production machine learning inference systems at scale.
  • Strong understanding of performance trade-offs in serving large AI models.
  • Experience designing reliable distributed systems or production infrastructure.
  • Understanding of GPU-backed AI workloads and scaling challenges.
  • Strong engineering fundamentals and ability to own complex technical problems.
  • Comfortable in a fast-moving environment with evolving systems.

Benefits

  • Competitive base pay for the Bellevue market.
  • Eligibility for merit increases, annual bonuses, and stock based on performance.
  • Access to medical, dental, and vision insurance.
  • 401(k) plan with company match.
  • Paid holidays each calendar year.

Tech Stack

Categories

AI & MLBackendData Engineering