
Forward Deployed Engineer (Inference & Post-Training) - Mandarin Speaking
Together AI2 months ago
Singapore, SingaporeSenior
Responsibilities
- Select, configure, and optimize inference engines based on hardware, model architecture, and workload profile.
- Develop configuration updates and tune KV cache, speculative decoding, tensor parallelism, and quantization to meet throughput and latency targets.
- Drive hands-on reinforcement-learning training runs and optimize system design.
- Guide customers through LoRA, SFT, DPO, RLHF, and GRPO pipelines from experimentation through production.
- Serve as the primary technical contact for strategic accounts, monitoring endpoint configurations and optimizing customer deployments.
- Establish inference and post-training alignment during onboarding to improve customer time-to-value.
- Surface field insights to influence the software and model roadmap and contribute product changes where needed.
- Drive early feature and research adoption with strategic customers.
Requirements
- 5+ years of technical experience focused on inference systems, open-source LLM deployment, or post-training workflows.
- Expert hands-on experience with inference engines such as vLLM, TensorRT-LLM, or SGLang, including diagnosing engine-level performance issues.
- Deep knowledge of KV cache tuning, speculative decoding, tensor parallelism, pipeline parallelism, and quantization.
- Hands-on experience with fine-tuning and post-training pipelines including LoRA, SFT, DPO, RLHF, and GRPO.
- Broad knowledge of state-of-the-art open-source models and the ability to select models for customer use cases, hardware, and performance targets.
- Strong Python skills and comfort working in production environments.
- Mandarin-speaking ability is indicated by the job title.
- Must be a permanent resident or citizen of Singapore.
Benefits
- Startup equity, health insurance, and other benefits are offered.
- Flexible remote work arrangements are available.
- Applicants must be permanent residents or citizens of Singapore.
Tech Stack
Categories
Forward DeployedML Engineering
About Together AI
Together AI builds an AI-native cloud platform for developers, offering high-performance inference, fine-tuning/model shaping, and large-scale pre-training on on-demand GPU clusters with APIs and managed services. It emphasizes open-source models that teams can run and adapt, and also provides infrastructure for decentralized and scalable workloads. Founded in 2022 and headquartered in San Francisco, it is privately held and reports notable customers including Cursor, ElevenLabs, Salesforce, and Zoom.