Crusoe

Principal Engineer, CAPE

Crusoe
Apply
2 months ago

Base Salary

$285k - $335k/yr

Responsibilities

  • Own the unified observability plane correlating GPU, networking, storage, orchestration, and workload signals.
  • Design a fleet-wide logical computer with unified health modeling, scheduling, and state management across tens of thousands of accelerators.
  • Build closed-loop autonomy that can diagnose, decide, and remediate live jobs without human intervention.
  • Optimize goodput through scheduling, placement, and maintenance decisions.
  • Develop predictive failure detection for GPU, NVLink, optics, and thermal degradation and enable proactive workload migration.
  • Detect stragglers and silent failures in distributed jobs and identify affected ranks, GPUs, and nodes.
  • Incorporate real-time energy availability, cost, and thermal headroom into workload scheduling and placement.
  • Create a digital twin to test failures, scheduling policies, and remediation logic before production deployment.
  • Develop guarded, auditable agentic operations that can propose and execute fixes.
  • Build zero-trust, identity-scoped, policy-checked, auditable multi-tenancy workflows.
  • Automate hardware burn-in and qualification before repaired or new nodes receive customer workloads.

Requirements

  • 10+ years building infrastructure-layer systems at scale, such as fleet management, distributed control planes, scheduler internals, or hardware lifecycle automation.
  • Deep experience designing distributed systems involving consensus, state reconciliation, closed-loop automation, and autonomous decisions against live production infrastructure.
  • Hands-on fluency with GPU/HPC infrastructure, including GPU health telemetry, NVLink, InfiniBand or RoCE fabrics, and hardware-level thermal and power behavior.
  • Track record designing and shipping large-scale observability or telemetry platforms correlating compute, network, and storage signals.
  • Experience defining architecture and standards for ambiguous 0-to-1 systems.
  • Strong software engineering fundamentals in at least one systems language, such as Go, Rust, or C++.
  • Experience applying ML or statistical methods to noisy operational telemetry, including failure prediction or anomaly detection, is preferred.
  • Exposure to zero-trust or policy-based multi-tenancy architectures is preferred but not required.

Benefits

  • Competitive compensation and equity packages, including Restricted Stock Units.
  • Paid time off, holidays, and leave programs.
  • Comprehensive health, dental, and vision insurance, plus HSA contributions.
  • Paid parental leave, life insurance, and short- and long-term disability coverage.
  • Professional development, tuition reimbursement, and mental health and wellness support.
  • Commuter benefits, cell phone stipend, 401(k) with company match up to 4% of salary, and volunteer time off.
  • Global travel insurance, emergency assistance, daily meal allowance, and location-specific perks.

Tech Stack

Categories

Crusoe

About Crusoe

1,001-5,000 employees

As the AI factory company, Crusoe is on a mission to accelerate the abundance of energy and intelligence. The company provides a reliable, scalable, cost-effective, and energy-first solution for AI infrastructure. By harnessing large-scale energy resources, building AI-optimized data centers, and delivering an AI cloud platform, Crusoe empowers its customers to build the future faster.

Contact me