GrepJob
Together AI

Forward Deployed Engineer (Inference & Post-Training) - Mandarin Speaking

Together AI
Apply
about 3 hours ago
Singapore, SingaporeSenior
H1B Sponsor

Responsibilities

  • Select, configure, and optimize inference engines based on various factors.
  • Develop configuration updates to enhance POCs and optimize customer deployments.
  • Drive hands-on RL training runs and optimize system design for customers.
  • Act as the primary technical point of contact for strategic accounts.
  • Establish alignment with customers during onboarding for optimal configurations.
  • Influence product and model roadmaps by providing field insights.

Requirements

  • 5+ years of experience in a technical role focused on inference systems.
  • Expert-level experience with inference engines like vLLM and TensorRT-LLM.
  • Deep knowledge of KV cache tuning, speculative decoding, and quantization techniques.
  • Hands-on experience with fine-tuning and post-training pipelines.
  • Broad knowledge of state-of-the-art open-source models.
  • Strong Python skills and comfort in production environments.

Benefits

  • Competitive compensation and startup equity.
  • Health insurance and other benefits.
  • Flexibility in remote work arrangements.

Tech Stack

Categories

AI & MLData Science