9 hours ago
Base Salary
$270k - $402k/yr
Responsibilities
- Own technical execution for customer engagements from discovery and architecture through proofs of concept, deployment, cluster bring-up, validation, production readiness, and operational handoff.
- Translate ambiguous customer requirements into architectures, implementation plans, test criteria, runbooks, and engineering actions.
- Work hands-on across Linux, bare-metal infrastructure, Kubernetes, Slurm, networking, storage, observability, automation, and Groq platform integrations.
- Support GPU and LPX deployments through infrastructure bring-up, cluster validation, workload testing, benchmarking, failure isolation, and acceptance activities.
- Analyze AI training and inference workloads, including concurrency, throughput, latency, data movement, caching, scheduling, and infrastructure bottlenecks.
- Lead customer discovery, architecture reviews, demonstrations, and proofs of concept while explaining tradeoffs, results, and risks.
- Partner with Networking and Security teams on connectivity, routing, load balancing, access controls, security architecture, and enterprise diligence.
- Troubleshoot cross-functional production and pre-production issues and drive them to resolution.
- Build reusable tooling, automation, reference architectures, test suites, deployment patterns, documentation, and platform improvements.
- Bring recurring customer feedback and field evidence back to Product and Engineering.
Requirements
- At least 4 years of hands-on experience building, deploying, operating, or troubleshooting cloud infrastructure, AI infrastructure, HPC systems, large-scale platforms, or comparable production environments.
- Strong Linux and distributed-systems fundamentals with practical experience in Kubernetes, Slurm, bare-metal environments, or comparable infrastructure platforms.
- Technical depth in GPU or accelerator systems, networking, storage, orchestration/platform engineering, or infrastructure reliability, with breadth across adjacent layers.
- Working knowledge of AI training and inference workloads and their effects on compute, networking, storage, scheduling, latency, and throughput.
- Strong Python, Go, Bash, or equivalent scripting and programming skills for diagnostics, automation, deployment tooling, testing, or integrations.
- Demonstrated experience personally debugging and delivering systems rather than working only at the architecture, project-management, or escalation level.
- Ability to turn ambiguous problems into concrete technical actions and drive resolution across multiple teams.
- Strong written and verbal communication skills for customer requirements gathering, technical explanations, and documentation.
- Preferred experience with neoclouds, hyperscalers, AI infrastructure providers, HPC, frontier AI, or large-scale accelerator infrastructure.
- Preferred hands-on experience with NVIDIA GPU infrastructure, CUDA, NCCL, NVLink/NVSwitch, DCGM, GPU Operator, InfiniBand, RoCE, Kubernetes, or Slurm.
- Preferred experience operating multi-node GPU clusters, including health checks, burn-in or stress testing, collective-communication testing, benchmarking, and acceptance criteria.
- Preferred familiarity with VAST, Weka, Lustre, Ceph, or similar high-performance storage systems.
- Preferred experience with Terraform, Ansible, CI/CD, BMC/Redfish, PXE/iPXE, or related infrastructure lifecycle tooling.
- Preferred solutions engineering, sales engineering, solutions architecture, technical pre-sales, networking, security, or data center systems experience.
- Preferred customer-facing networking and security experience, including private interconnects, peering, BGP, routing, load balancing, Cilium network policy, access-control architecture, network isolation, and enterprise security reviews.
- Preferred experience defining or executing technical proofs of concept, reference architectures, cluster acceptance tests, performance benchmarks, migration plans, or production-readiness criteria.
Benefits
- Groq offers a Long-Term Incentive Program and a robust suite of employee benefits.
- Hiring is prioritized in or near the San Francisco Bay Area, New York City, and Dallas; compensation varies by location and international market.
Tech Stack
Categories
Forward Deployed
About Groq
Groq designs and operates AI inference hardware and cloud services built around its LPU architecture, aimed at developers and enterprises running large language models and other ML workloads. It generates revenue by selling processors and offering pay-as-you-go inference through its hosted platform and APIs. Founded in 2016 and headquartered in Mountain View, California, the company is privately held.
