
Lead AI Platform
Integrant, Inc.5 months ago
Cairo, EgyptStaff+
Responsibilities
- Translate AI/ML workloads into optimized infrastructure and deployment strategies.
- Optimize model latency, throughput, memory utilization, and GPU performance across environments.
- Design and implement training and inference pipelines using TensorRT, Triton, and NIM.
- Convert and optimize models across PyTorch, ONNX, and TensorRT.
- Analyze compute, memory, and network bottlenecks using GPU and system profiling tools.
- Improve GPU utilization and scheduling efficiency across clusters.
- Design scalable distributed training and inference architectures for multi-GPU and multi-node systems.
- Work with customers to define AI infrastructure strategies and deployment models.
- Support production deployments, including monitoring, rollback, and performance validation.
- Conduct applied research on model efficiency and infrastructure utilization.
- Mentor team members on AI infrastructure, optimization, and GPU systems.
- Use experiment tracking tools to compare parameters, metrics, and artifacts.
- Investigate post-deployment model degradation, including concept drift, data pipeline changes, and traffic pattern shifts.
- Perform root cause analysis for ML systems by isolating variables and reproducing issues.
Requirements
- 8+ years of experience in AI systems and 8+ years of experience in ML systems, HPC, and AI infrastructure.
- Strong proficiency in Python.
- Strong experience with GPU-based AI workloads and performance optimization.
- Deep understanding of quantization, pruning, batching, and other model optimization techniques.
- Hands-on experience with PyTorch, ONNX, ONNX Runtime, TensorRT, TensorRT-LLM, and Triton Inference Server.
- Knowledge of CUDA, cuDNN, and GPU architecture fundamentals.
- Experience with distributed multi-GPU and multi-node systems.
- Familiarity with NCCL, NVLink, InfiniBand, Kubernetes, or Slurm.
- Experience deploying AI models into production environments.
- Experience analyzing system bottlenecks and using profiling tools such as Nsight and TensorRT profiler.
- Knowledge of GPU workload cost optimization strategies.
- Experience with MLflow, W&B, or Neptune for experiment tracking.
- Nice to have experience with NVIDIA NIM, the NGC ecosystem, Megatron-LM, NeMo, large-scale LLM training or inference, and LLM optimization techniques such as KV cache and batching strategies.
- Nice to have familiarity with MLOps practices, CI/CD for AI systems, customer-facing architecture or consulting, and hybrid cloud or on-premises HPC environments.
Benefits
- Hybrid workplace in Cairo, Egypt.
- Full-time employment.
- Salary paid in USD.
- Six-month career-advancing opportunities.
- Supportive and friendly work environment.
- Premium medical insurance for employees and families.
- English language development courses.
- Interest-free loans paid over 2.5 years.
- Technical development courses.
- Planned overtime program.
- Employment referral program.
- Premium location in Maadi.
- Social insurance.