7 hours ago
Paris, FranceSenior
Responsibilities
- Build zonal-resilience automation for safe workload evacuations, switchovers, and recovery.
- Design and build infrastructure- and application-level fault-injection systems, incident replay, and controlled resilience experiments.
- Develop blast-radius controls, kill switches, validation mechanisms, and rollback paths for contained production experiments.
- Build agents and automation that propose failure scenarios, triage experiment results, and connect reliability findings to remediation and verification.
- Lead gamedays from hypothesis and scenario design through execution, documentation, remediation tracking, and validation.
- Design and implement distributed systems, including gRPC services, Kubernetes controllers, and shared platform components.
- Contribute to technical designs, platform strategy, engineering practices, and mentoring.
Requirements
- Strong fundamentals in distributed systems, including consistency, failure modes, backpressure, idempotency, quorum, retries, and failure recovery.
- Understanding of Kubernetes workload lifecycles, including pods, controllers, scheduling, draining, and eviction.
- Experience designing, building, or operating production systems where safety, availability, and controlled failure handling are important.
- Ability to communicate technical decisions through design documents, runbooks, postmortems, and cross-functional discussions.
- Ability to collaborate across engineering teams to understand unfamiliar systems, identify failure modes, and drive resilience improvements.
- Experience with reliability engineering, chaos engineering, zonal failover, AI-assisted operational workflows, traffic interception, or large-scale observability systems is beneficial but not required.
Benefits
- Hybrid workplace designed to support collaboration and work-life harmony.
- Develop expertise in distributed systems, production resilience, Kubernetes, and large-scale infrastructure.
- Work on reliability systems operating across Datadog’s production environment.
- Gain experience with fault injection, zonal resilience, incident reproduction, and AI-assisted reliability workflows.
- Collaborate with infrastructure, database, observability, and service teams.
- Mentor engineers and contribute to technical designs, engineering practices, and platform strategy.
- Benefits may vary based on country of employment and the nature of employment.
Tech Stack
gRPCKubernetes
Categories
Site Reliability
About Datadog
Datadog builds a SaaS observability and security platform that monitors infrastructure, applications, logs, and services for engineering and DevOps teams. Its products include infrastructure monitoring, APM, log management, real user monitoring, and cloud security, sold via subscriptions and used across cloud-native and hybrid environments. Founded in 2010 and headquartered in New York City, Datadog is a public company trading on NASDAQ under the ticker DDOG.
