
Forward Deployed Engineer (Inference & Post-Training) - Mandarin Speaking
Together AIabout 3 hours ago
Responsibilities
- Select, configure, and optimize inference engines based on various factors.
- Develop configuration updates to enhance POCs and optimize customer deployments.
- Drive hands-on RL training runs and optimize system design for customers.
- Act as the primary technical point of contact for strategic accounts.
- Establish alignment with customers during onboarding for optimal configurations.
- Influence product and model roadmaps by providing field insights.
Requirements
- 5+ years of experience in a technical role focused on inference systems.
- Expert-level experience with inference engines like vLLM and TensorRT-LLM.
- Deep knowledge of KV cache tuning, speculative decoding, and quantization techniques.
- Hands-on experience with fine-tuning and post-training pipelines.
- Broad knowledge of state-of-the-art open-source models.
- Strong Python skills and comfort in production environments.
Benefits
- Competitive compensation and startup equity.
- Health insurance and other benefits.
- Flexibility in remote work arrangements.