12 hours ago
Remote, SerbiaSenior
Responsibilities
- Build and improve the inference layer of the Gcore Inference platform.
- Integrate and operate inference frameworks including vLLM, SGLang, NVIDIA Dynamo, and TensorRT-LLM.
- Bring new language and multimodal models into production.
- Improve inference latency, throughput, memory usage, GPU utilization, and cost efficiency.
- Debug performance and reliability issues across model code, inference frameworks, GPU execution, networking, and Kubernetes.
- Collaborate with platform, infrastructure, product, and customer-facing teams to deliver reliable product features.
- Contribute improvements to open-source inference projects when appropriate.
Requirements
- 5+ years of experience writing reliable, well-tested production code.
- Strong Python skills and experience designing production systems.
- Hands-on experience with PyTorch and deploying machine learning models.
- Experience with Linux, Docker, and Kubernetes.
- Experience in at least one of distributed systems, GPU computing, ML runtimes, model optimization, or cluster scheduling.
- Ability to debug complex problems across software, infrastructure, and hardware.
- Strong developer-experience focus and effective communication and collaboration skills.
- Interest in inference engineering; direct experience with inference engines is not required.
- Preferred: experience with vLLM, SGLang, NVIDIA Dynamo, TensorRT-LLM, or a similar inference framework.
- Preferred: experience running GPU workloads in production and with inference techniques such as quantization, continuous batching, speculative decoding, prefix caching, chunked prefill, or LoRA serving.
- Preferred: experience with CUDA, Triton, TensorRT, or other GPU programming tools.
- Preferred: experience profiling and improving model latency, throughput, memory usage, or GPU utilization.
- Preferred: experience with distributed inference, multi-GPU systems, scheduling, or autoscaling.
- Preferred: contributions to open-source ML, inference, or infrastructure projects.
Benefits
- Competitive compensation.
- Flexible working hours and hybrid or remote options depending on the role.
- Work from anywhere in the world for up to 45 days per year.
- Private medical insurance for the employee and family, with benefits varying by location.
- Extra paid vacation and sick leave days, with benefits varying by location.
- Support for important life moments and celebrations.
- Language courses.
- Modern offices with snacks, drinks, and entertainment, where applicable.
- Team sports and social activities, where applicable.
- Position available only under an employment (labor) agreement.
Tech Stack
Categories
About Gcore
Gcore provides edge and cloud infrastructure for enterprises and service providers, including CDN, DNS, traffic management, security, and GPU cloud for AI workloads. It operates its own global network across six continents to deliver low-latency performance for media, gaming, and enterprise applications. Founded in 2014 and headquartered in Contern, Luxembourg, Gcore is privately held.
