6 months ago
Beijing, ChinaMid Level / Senior
Responsibilities
- Build and scale AI inference infrastructure, including inference serving, scheduling, orchestration, and autoscaling.
- Design and optimize CPU/GPU resource management systems for utilization and cost efficiency.
- Implement production GPU scheduling and multiplexing technologies, including MIG, MPS, and virtualization.
- Optimize inference throughput, latency, and stability for high-concurrency and complex business scenarios.
- Contribute to system reliability, disaster recovery, observability, and cloud cost governance.
- Explore AI-native infrastructure and automated operations and maintenance systems.
Requirements
- 3–5 years of experience in backend or infrastructure engineering, with cloud-native or AI platform experience preferred.
- Strong proficiency in Go or Python and solid software engineering fundamentals.
- Deep understanding of Linux internals, networking, and distributed systems.
- Hands-on experience with Kubernetes, Docker, and microservices architecture.
- Prior project experience in inference systems, task scheduling, or resource management.
- Preferred experience includes GPU platforms, customized Kubernetes schedulers, Ray, model-serving frameworks, distributed inference, MIG, MPS, vGPU, SRE, observability, or cloud cost optimization.
- Open-source contributions or significant technical influence in the engineering community are advantageous.
Benefits
- Competitive salary, equity, and benefits package.
- Flexible work environment with remote and on-site options.
- Unlimited, flexible time off.
- Stock options available for core team members.
- 401(k) plan.
- Comprehensive health, dental, and vision insurance.
- Latest office equipment.
- Opportunities for professional growth and development.
