19 days ago
Paris, FranceSenior
Responsibilities
- Analyze AI workloads to identify kernel-, runtime-, memory-, and system-level bottlenecks on Arago’s accelerator.
- Develop and optimize custom kernels, fused operators, and execution strategies to maximize device utilization.
- Design model and operator mappings across multiple devices, including communication and synchronization strategies.
- Develop inference-serving techniques including continuous batching, paged KV caches, prefix/context caching, chunked prefill, and prefill/decode interleaving or disaggregation.
- Build profiling, benchmarking, and performance-analysis infrastructure for kernels, full models, and serving workloads.
- Collaborate with hardware, compiler, and runtime teams to co-design software abstractions and influence future hardware features.
Requirements
- Strong experience in high-performance ML inference, GPU or accelerator programming, or ML systems engineering.
- Deep understanding of computer architecture, accelerator and GPU execution models, memory hierarchies, parallelism, and performance bottlenecks.
- Experience developing and optimizing custom kernels using CUDA, Triton, ROCm/HIP, or equivalent low-level programming environments.
- Experience with operator fusion, tiling, scheduling, data movement optimization, graph execution, and profiling compute- and memory-bound workloads.
- Strong understanding of distributed model execution, including tensor, pipeline, sequence, or expert parallelism and communication/computation overlap.
- Hands-on experience with inference-serving systems such as vLLM, SGLang, TensorRT-LLM, or equivalent, including KV-cache management, continuous batching, paged attention, and prefill/decode scheduling.
- Strong C++ and Python skills and comfort working with actively evolving compiler, runtime, kernel, and accelerator abstractions.
- Proficient English language skills.
- Exposure to or experience with MLIR and MLIR dialects is a strong plus.
Benefits
- Competitive cash compensation based on location, experience, and comparable team-member pay; no specific amount is stated.
- Meaningful stock option plan, included in the majority of full-time offers.
- Healthcare coverage with family-friendly options, pension contributions, professional development support, and 25 days of PTO in addition to public holidays.
- Ownership of a key technical domain with opportunities for vertical and horizontal growth based on performance and individual drive.
