12 months ago
Responsibilities
- Build and operate datasets, training and evaluation pipelines, benchmarks, and internal tooling.
- Implement models, run experiments at scale, and profile systems for reliability, performance, and cost.
- Orchestrate distributed training and distributed reinforcement learning with Ray, including scheduling, scaling, and failure recovery.
- Improve observability, reproducibility, and usability of the research stack.
- Establish automated benchmarks and regression tests for forecasting, anomaly detection, multimodal analysis, agents, and code repair.
- Collaborate with Research Scientists, Product, and Engineering to integrate AI capabilities into Datadog products and harden prototypes into reliable services.
- Contribute high-quality code, documentation, and open-source artifacts for reproducibility, extension, and evaluation.
Requirements
- Strong software engineering skills with experience in observability, SRE, or security domains.
- Depth in distributed computing and ML systems for training and inference at scale; experience with Ray, Slurm, or similar frameworks is a plus.
- Proficiency in Python and familiarity with a systems language such as Rust, C++, or Go.
- Comfort with modern cloud and data infrastructure.
- Practical experience implementing and operating ML training and inference systems with technologies such as PyTorch or JAX, including containerization, orchestration, and GPU acceleration.
- Familiarity with efficient training, fine-tuning, and inference techniques for large foundation models.
- Ability to explain design and performance trade-offs to technical and non-technical audiences.
- Strong interest in open science and open-source contributions, including rigorous benchmarks and shared artifacts.
- Bonus: experience bridging research prototypes and product applications, particularly with foundation models, generative AI agents, or domain-specific LLM deployments.
- Bonus: hands-on GPU programming and optimization, including CUDA.
- Bonus: experience writing production data pipelines and applications.
- Bonus: experience supporting or contributing to research publications.
Benefits
- Competitive global benefits, new-hire stock equity through RSUs, and an employee stock purchase plan.
- Opportunity to collaborate with colleagues across Datadog offices in New York City and Paris.
- Opportunities to attend and present at conferences and meetups.
- Intra-departmental mentor and buddy program.
- Inclusive company culture and access to Datadog Community Guilds.
- Benefits may vary by country and employment arrangement.
About Datadog
Datadog is the essential monitoring platform for cloud applications. We bring together data from servers, containers, databases, and third-party services to make your stack entirely observable. These capabilities help DevOps teams avoid downtime, resolve performance issues, and ensure customers are getting the best user experience.