Integrant, Inc.

Lead AI Platform

Integrant, Inc.
Apply
5 months ago
Cairo, EgyptStaff+

Responsibilities

  • Translate AI/ML workloads into optimized infrastructure and deployment strategies.
  • Optimize model latency, throughput, memory utilization, and GPU performance across environments.
  • Design and implement training and inference pipelines using TensorRT, Triton, and NIM.
  • Convert and optimize models across PyTorch, ONNX, and TensorRT.
  • Analyze compute, memory, and network bottlenecks using GPU and system profiling tools.
  • Improve GPU utilization and scheduling efficiency across clusters.
  • Design scalable distributed training and inference architectures for multi-GPU and multi-node systems.
  • Work with customers to define AI infrastructure strategies and deployment models.
  • Support production deployments, including monitoring, rollback, and performance validation.
  • Conduct applied research on model efficiency and infrastructure utilization.
  • Mentor team members on AI infrastructure, optimization, and GPU systems.
  • Use experiment tracking tools to compare parameters, metrics, and artifacts.
  • Investigate post-deployment model degradation, including concept drift, data pipeline changes, and traffic pattern shifts.
  • Perform root cause analysis for ML systems by isolating variables and reproducing issues.

Requirements

  • 8+ years of experience in AI systems and 8+ years of experience in ML systems, HPC, and AI infrastructure.
  • Strong proficiency in Python.
  • Strong experience with GPU-based AI workloads and performance optimization.
  • Deep understanding of quantization, pruning, batching, and other model optimization techniques.
  • Hands-on experience with PyTorch, ONNX, ONNX Runtime, TensorRT, TensorRT-LLM, and Triton Inference Server.
  • Knowledge of CUDA, cuDNN, and GPU architecture fundamentals.
  • Experience with distributed multi-GPU and multi-node systems.
  • Familiarity with NCCL, NVLink, InfiniBand, Kubernetes, or Slurm.
  • Experience deploying AI models into production environments.
  • Experience analyzing system bottlenecks and using profiling tools such as Nsight and TensorRT profiler.
  • Knowledge of GPU workload cost optimization strategies.
  • Experience with MLflow, W&B, or Neptune for experiment tracking.
  • Nice to have experience with NVIDIA NIM, the NGC ecosystem, Megatron-LM, NeMo, large-scale LLM training or inference, and LLM optimization techniques such as KV cache and batching strategies.
  • Nice to have familiarity with MLOps practices, CI/CD for AI systems, customer-facing architecture or consulting, and hybrid cloud or on-premises HPC environments.

Benefits

  • Hybrid workplace in Cairo, Egypt.
  • Full-time employment.
  • Salary paid in USD.
  • Six-month career-advancing opportunities.
  • Supportive and friendly work environment.
  • Premium medical insurance for employees and families.
  • English language development courses.
  • Interest-free loans paid over 2.5 years.
  • Technical development courses.
  • Planned overtime program.
  • Employment referral program.
  • Premium location in Maadi.
  • Social insurance.

Tech Stack

Integrant, Inc.

About Integrant, Inc.

201-500 employees
Contact me