Responsibilities
- Own and operate GPU and accelerator clusters for model training, inference, evaluation, and experimentation.
- Design local GPU infrastructure and compute systems that reduce external-provider dependency and control costs.
- Build scheduling, orchestration, placement, quota management, workload isolation, and utilization systems across heterogeneous accelerators.
- Optimize inference pipelines for latency, throughput, reliability, memory efficiency, and cost using serving frameworks such as vLLM, Triton Inference Server, and TensorRT.
- Partner with ML engineers and researchers to improve training, evaluation, batch inference, online inference, deployment, and production-debugging workflows.
- Build observability for GPU utilization, memory pressure, queue depth, saturation, token throughput, request latency, failed workloads, capacity pressure, and spend.
- Drive reliability engineering, incident response, alerting, runbooks, and post-incident improvements for always-on AI compute infrastructure.
- Evaluate and integrate new hardware, cloud instance families, specialized accelerators, runtimes, schedulers, and serving frameworks.
- Build internal tooling that makes GPU usage visible, accountable, and easier for teams to consume.
- Contribute to architecture decisions balancing performance, cost efficiency, scalability, operational simplicity, and production safety.
Requirements
- 5+ years of infrastructure engineering experience, including significant experience with GPU compute, ML infrastructure, distributed systems, high-performance computing, or large-scale production platforms.
- Hands-on experience operating GPU clusters or accelerator-backed infrastructure in production or production-like environments, including scheduling, orchestration, utilization monitoring, and cost optimization.
- Strong systems engineering fundamentals across Linux, networking, storage, containers, Kubernetes, distributed runtimes, and production debugging.
- Experience with ML serving frameworks such as vLLM, Triton Inference Server, TensorRT, TorchServe, KServe, Ray Serve, or equivalent systems.
- Proficiency in Python for infrastructure automation, tooling, debugging, integration, and operational workflows.
- Understanding of performance tradeoffs involving batching, concurrency, memory usage, GPU utilization, model size, latency, throughput, availability, and cost.
- Track record of optimizing compute costs while maintaining performance, reliability, and availability expectations.
- Experience building observable systems with metrics, logs, traces, dashboards, alerts, and incident workflows.
- Ability to work in high-stakes, always-on environments where uptime, throughput, correctness, and operational discipline are critical.
- Clear communication skills for translating infrastructure tradeoffs to researchers, product teams, platform engineers, security stakeholders, and engineering leadership.
- Preferred experience includes work at a frontier AI lab, hyperscaler, high-frequency trading firm, research platform, or high-scale ML organization.
- Preferred qualifications include familiarity with custom silicon or specialized accelerators, capacity planning, procurement input, reserved capacity strategy, cloud accelerator economics, or GPU fleet cost management.
- Preferred experience includes distributed training frameworks such as DeepSpeed, Megatron-LM, FSDP, or Ray; debugging CUDA, NCCL, kernels, drivers, runtimes, memory, or networking; and using Rust, C++, Go, or CUDA.
- Crypto, financial services, trading infrastructure, or security-sensitive production infrastructure experience is a plus.
Categories
About Kraken
Kraken helps people own the power of their money. Founded in 2011, we're one of the world's longest-standing crypto platforms, now offering crypto, stocks, futures, staking, payments and more. Trusted by millions of institutions, professional traders, and consumers across 190+ countries. Built on the belief that everyone should own the power of their money, Kraken has been a category leader in crypto since 2011. We've grown into a global financial platform offering crypto, stocks, futures, staking, payments and institutional infrastructure — all on a single foundation built for the future of finance. Our product lineup: — Kraken: the world's leading hybrid investment platform, built for the future of finance from the ground up — Kraken Pro: the all-in command center that unleashes your trading potential — Krak: the next-generation money platform that turns every type of asset into usable money Proof of Reserves since 2014. Assets held 1:1. Trusted by 13M+ people, thousands of institutions, and a growing community of developers building on Kraken rails. Our mission: accelerate the global adoption of crypto so everyone can achieve financial freedom and inclusion. Risk Disclosure: https://www.kraken.com/en-ca/social-media-disclosure
