Kraken

Senior AI Compute Infrastructure Engineer

Kraken
Apply
4 months ago
San José, Costa Rica +2 moreSenior

Responsibilities

  • Own and operate GPU and accelerator clusters for model training, inference, evaluation, and experimentation.
  • Design local GPU infrastructure and compute systems that reduce external-provider dependency and control costs.
  • Build scheduling, orchestration, placement, quota management, workload isolation, and utilization systems across heterogeneous accelerators.
  • Optimize inference pipelines for latency, throughput, reliability, memory efficiency, and cost using serving frameworks such as vLLM, Triton Inference Server, and TensorRT.
  • Partner with ML engineers and researchers to improve training, evaluation, batch inference, online inference, deployment, and production-debugging workflows.
  • Build observability for GPU utilization, memory pressure, queue depth, saturation, token throughput, request latency, failed workloads, capacity pressure, and spend.
  • Drive reliability engineering, incident response, alerting, runbooks, and post-incident improvements for always-on AI compute infrastructure.
  • Evaluate and integrate new hardware, cloud instance families, specialized accelerators, runtimes, schedulers, and serving frameworks.
  • Build internal tooling that makes GPU usage visible, accountable, and easier for teams to consume.
  • Contribute to architecture decisions balancing performance, cost efficiency, scalability, operational simplicity, and production safety.

Requirements

  • 5+ years of infrastructure engineering experience, including significant experience with GPU compute, ML infrastructure, distributed systems, high-performance computing, or large-scale production platforms.
  • Hands-on experience operating GPU clusters or accelerator-backed infrastructure in production or production-like environments, including scheduling, orchestration, utilization monitoring, and cost optimization.
  • Strong systems engineering fundamentals across Linux, networking, storage, containers, Kubernetes, distributed runtimes, and production debugging.
  • Experience with ML serving frameworks such as vLLM, Triton Inference Server, TensorRT, TorchServe, KServe, Ray Serve, or equivalent systems.
  • Proficiency in Python for infrastructure automation, tooling, debugging, integration, and operational workflows.
  • Understanding of performance tradeoffs involving batching, concurrency, memory usage, GPU utilization, model size, latency, throughput, availability, and cost.
  • Track record of optimizing compute costs while maintaining performance, reliability, and availability expectations.
  • Experience building observable systems with metrics, logs, traces, dashboards, alerts, and incident workflows.
  • Ability to work in high-stakes, always-on environments where uptime, throughput, correctness, and operational discipline are critical.
  • Clear communication skills for translating infrastructure tradeoffs to researchers, product teams, platform engineers, security stakeholders, and engineering leadership.
  • Preferred experience includes work at a frontier AI lab, hyperscaler, high-frequency trading firm, research platform, or high-scale ML organization.
  • Preferred qualifications include familiarity with custom silicon or specialized accelerators, capacity planning, procurement input, reserved capacity strategy, cloud accelerator economics, or GPU fleet cost management.
  • Preferred experience includes distributed training frameworks such as DeepSpeed, Megatron-LM, FSDP, or Ray; debugging CUDA, NCCL, kernels, drivers, runtimes, memory, or networking; and using Rust, C++, Go, or CUDA.
  • Crypto, financial services, trading infrastructure, or security-sensitive production infrastructure experience is a plus.
Kraken

About Kraken

1,001-5,000 employees

Kraken helps people own the power of their money. Founded in 2011, we're one of the world's longest-standing crypto platforms, now offering crypto, stocks, futures, staking, payments and more. Trusted by millions of institutions, professional traders, and consumers across 190+ countries. Built on the belief that everyone should own the power of their money, Kraken has been a category leader in crypto since 2011. We've grown into a global financial platform offering crypto, stocks, futures, staking, payments and institutional infrastructure — all on a single foundation built for the future of finance. Our product lineup: — Kraken: the world's leading hybrid investment platform, built for the future of finance from the ground up — Kraken Pro: the all-in command center that unleashes your trading potential — Krak: the next-generation money platform that turns every type of asset into usable money Proof of Reserves since 2014. Assets held 1:1. Trusted by 13M+ people, thousands of institutions, and a growing community of developers building on Kraken rails. Our mission: accelerate the global adoption of crypto so everyone can achieve financial freedom and inclusion. Risk Disclosure: https://www.kraken.com/en-ca/social-media-disclosure

Contact me