15 days ago
Base Salary
$270k - $402k/yr
Responsibilities
- Own technical execution across customer engagements from discovery and architecture through proofs of concept, deployment, validation, production readiness, and handoff.
- Translate ambiguous customer requirements into architectures, implementation plans, test criteria, runbooks, and engineering actions.
- Bring up and troubleshoot Linux, bare-metal, Kubernetes, Slurm, GPU, LPX, networking, storage, observability, automation, and Groq platform environments.
- Support large-scale GPU and LPX deployments through cluster validation, workload testing, benchmarking, failure isolation, and acceptance testing.
- Reason about AI training and inference workloads, including concurrency, throughput, latency, data movement, caching, scheduling, and infrastructure bottlenecks.
- Lead customer discovery, architecture reviews, demonstrations, and proofs of concept while explaining technical tradeoffs and risks.
- Partner with networking and security teams on connectivity, peering, routing, load balancing, access controls, security reviews, and enterprise diligence.
- Troubleshoot cross-functional production and pre-production issues and drive them through resolution.
- Build reusable tools, automation, reference architectures, test suites, deployment patterns, and documentation.
- Bring customer feedback and field evidence to Product and Engineering to improve repeatable platform capabilities.
Requirements
- At least 4 years of hands-on experience building, deploying, operating, or troubleshooting cloud infrastructure, AI infrastructure, HPC systems, large-scale platforms, or comparable production environments.
- Strong Linux and distributed-systems fundamentals with practical experience in Kubernetes, Slurm, bare-metal environments, or comparable platforms.
- Technical depth in GPU or accelerator systems, networking, storage, orchestration/platform engineering, or infrastructure reliability, with breadth across adjacent layers.
- Working knowledge of AI training and inference workloads and their effects on compute, networking, storage, scheduling, latency, and throughput.
- Strong Python, Go, Bash, or equivalent scripting and programming skills.
- Demonstrated experience personally debugging and delivering systems rather than working only at the architecture, project-management, or escalation level.
- Ability to decompose ambiguous problems, drive resolution across multiple teams, and communicate technical tradeoffs clearly.
- Preferred experience with neoclouds, hyperscalers, AI infrastructure providers, HPC environments, frontier AI companies, or large-scale accelerator infrastructure.
- Preferred hands-on experience with NVIDIA GPU infrastructure, CUDA, NCCL, NVLink/NVSwitch, DCGM, GPU Operator, InfiniBand, RoCE, Kubernetes, or Slurm.
- Preferred experience with multi-node GPU clusters, health checks, stress testing, collective-communication testing, performance benchmarking, and acceptance criteria.
- Preferred familiarity with VAST, Weka, Lustre, Ceph, or similar high-performance storage systems.
- Preferred experience with Terraform, Ansible, CI/CD, BMC/Redfish, PXE/iPXE, or related lifecycle tooling.
- Preferred customer-facing solutions engineering, sales engineering, solutions architecture, technical pre-sales, networking, security, or data center systems experience.
- Preferred experience with private interconnects, peering, BGP, routing, load balancing, Kubernetes/Cilium network policy, traffic design, access-control architecture, network isolation, SOC 2, or ISO 27001-style diligence.
Benefits
- The posting mentions a Long-Term Incentive Program and a robust suite of employee benefits.
- Hiring is prioritized in or near the San Francisco Bay Area, New York City, and Dallas.
- Compensation and benefits vary for international candidates based on local market dynamics.
Tech Stack
Categories
Forward Deployed
About Groq
Groq designs and operates AI inference hardware and cloud services built around its LPU architecture, aimed at developers and enterprises running large language models and other ML workloads. It generates revenue by selling processors and offering pay-as-you-go inference through its hosted platform and APIs. Founded in 2016 and headquartered in Mountain View, California, the company is privately held.
