
Devops & SysOps Architect
Integrant, Inc.4 months ago
Cairo, EgyptStaff+
Responsibilities
- Partner with sales and solution teams to identify and qualify opportunities and support technical presales activities.
- Lead discovery workshops, RFP responses, architecture presentations, proof-of-concepts, and client technical discussions.
- Operate, troubleshoot, and optimize production Kubernetes clusters and GPU/HPC environments in client accounts.
- Own deep Linux administration, including kernel tuning, storage, networking, and performance profiling.
- Implement and maintain infrastructure-as-code pipelines, GitOps workflows, and CI/CD systems.
- Design end-to-end cloud, hybrid, and on-premises HPC platform architectures, including workload isolation, networking, and storage.
- Design and operate GPU compute platforms, including GPU Operator deployment, MIG partitioning, scheduling, and AI infrastructure.
- Define observability, SLO/SLA, alerting, incident response, disaster recovery, security, and multi-tenancy frameworks.
- Lead critical incident response and root-cause analysis and maintain runbooks, operational playbooks, and knowledge-base content.
- Recommend and validate technology choices and produce architecture decision records, solution blueprints, and technical runbooks.
Requirements
- 10+ years of platform or infrastructure engineering experience, including at least two years in an architect-level role.
- Production experience operating Kubernetes at scale in multi-cluster and multi-tenant environments.
- Significant low-level Linux systems administration experience covering kernel, networking, and storage.
- HPC and/or GPU infrastructure experience, including physical GPU servers, NCCL, InfiniBand, or high-speed fabrics.
- Demonstrable presales or client-facing experience.
- Production experience with Terraform and/or Ansible.
- Strong understanding of GitOps and enterprise CI/CD pipelines.
- Preferred experience with NVIDIA GPU Operator, MIG partitioning, Run:AI, or equivalent GPU scheduling tools.
- Preferred knowledge of distributed AI training infrastructure such as PyTorch DDP, Horovod, and DeepSpeed.
- Preferred familiarity with NVIDIA Triton Inference Server, TensorRT, Weka, Ceph, GPUDirect Storage, Vault, External Secrets, and zero-trust network architectures.
- Preferred exposure to bare-metal provisioning and HPC cluster management using Slurm, PBS, or equivalent.
- Advantageous certifications include CKA, CKS, RHCE, RHCA, AWS Solutions Architect, Azure Solutions Architect Expert, HashiCorp Terraform Associate, Vault Associate, and NVIDIA DLI certifications.
Benefits
- Competitive compensation package.
- PTO and full medical and dental coverage.
- Opportunity to travel and work onsite with U.S. customers.
- In-house technical and English training programs.
- Dedicated learning time through the 4Plus1 Program.
- Interest-free loans.
- Flexible work schedules.
- Workplace is hybrid.
- Additional perks include events, sponsored lunch, a game area, and rooftop hangout.
Tech Stack
Categories
DevOpsSolutions Engineering