Datadog

Staff Engineer, Compute

Datadog
Apply
4 hours ago
Paris, France +2 moreStaff+
H1B Sponsor

Responsibilities

  • Lead the technical direction of capacity management and workload placement for Datadog’s Kubernetes platform spanning more than 100,000 virtual machines.
  • Design and build systems that schedule and deploy engineering workloads across regions while balancing capacity constraints, reliability, and performance.
  • Partner with infrastructure teams to evolve multi-region and multi-cloud capacity orchestration.
  • Develop production software in Go to improve Kubernetes platform capabilities, automation, and operational efficiency.
  • Use data and capacity signals to influence infrastructure decisions, forecast growth, and improve workload placement strategies.

Requirements

  • Significant experience designing and operating large-scale Kubernetes-based infrastructure or platform systems.
  • Strong software engineering and programming skills, ideally in Go or a comparable systems programming language.
  • Hands-on experience with at least one major cloud provider: AWS, Google Cloud, or Azure.
  • Understanding of distributed cloud infrastructure and a strong systems mindset.
  • Experience with operational data, capacity forecasting, or analytical approaches that inform engineering decisions.
  • Proven technical leadership across teams and ability to influence architecture without direct management authority.

Benefits

  • Hybrid workplace arrangement.
  • Generous and competitive benefits package.
  • New hire stock equity (RSUs) and employee stock purchase plan.
  • Continuous career development and pathing opportunities.
  • Employee-focused onboarding.
  • Internal mentor and cross-departmental buddy program.
  • Friendly and inclusive workplace culture.
  • Benefits may vary by country of employment and nature of employment.

Tech Stack

Categories

Datadog

About Datadog

1,001-5,000 employees

Datadog is the essential monitoring platform for cloud applications. We bring together data from servers, containers, databases, and third-party services to make your stack entirely observable. These capabilities help DevOps teams avoid downtime, resolve performance issues, and ensure customers are getting the best user experience.