Principal Solutions Architect, AI Data Infrastructure
Firmus Technologies17 days ago
Singapore, SingaporeStaff+
Responsibilities
- Own the end-to-end AI data-path reference architecture across parallel and software-defined storage, NVMe and NVMe-oF, object and file systems, memory disaggregation and pooling, KV-cache, and data movement.
- Evaluate, benchmark, and select technologies against real AI workloads, translating results into performance, capacity, power-efficiency, and total-cost-of-ownership decisions.
- Advise customers by sizing solutions, validating workload requirements, and resolving production performance issues.
- Shape the internal platform roadmap and integrate storage and memory into Kubernetes-based and bare-metal environments with compute and networking teams.
- Apply RDMA, RoCEv2, and InfiniBand expertise to connect storage, memory, and GPUs.
- Establish operational practices for data protection, resilience, and lifecycle management.
- Integrate storage and memory into GPU clusters and assess CAPEX, server-count, and power trade-offs.
- Advise C-level customers, author reference architectures and whitepapers, and mentor engineers.
Requirements
- Typically 10+ years of experience in storage, memory, HPC, or AI infrastructure roles.
- Deep, vendor-agnostic expertise across storage, memory, caching, and the AI data path.
- Hands-on experience with high-performance storage and memory platforms such as WEKA, VAST Data, DDN, or Dell PowerScale/PowerFlex.
- Fluency in NVMe, NVMe-oF, parallel and software-defined storage, object and file systems, RDMA, RoCEv2, and InfiniBand.
- Knowledge of CXL, memory disaggregation, and KV-cache optimization for LLM workloads.
- Ability to design and run workload-driven evaluations across bare-metal, virtualized, and containerized environments.
- Experience integrating storage and memory into GPU clusters such as DGX, HGX, or similar systems.
- Ability to translate benchmarking results into architecture and business cases, including CAPEX, server-count, power, and TCO trade-offs.
- Strong communication skills for advising C-level customers, writing technical architectures or whitepapers, mentoring engineers, and operating in ambiguous environments.
Benefits
- Full-time employment based in Singapore.
- Opportunity to work on sustainable AI infrastructure, Firmus AI Cloud, and large-scale AI Factory systems.
Tech Stack
Categories
Solutions Engineering