about 4 hours ago
Responsibilities
- Lead strategic engineering initiatives for teams building, operating, and evolving services on the Observability Data Platform.
- Design and build scalable platform capabilities for cloud cost allocation, trend analysis, and optimization recommendations.
- Develop operational intelligence and self-service tooling for platform health, incident troubleshooting, and operational efficiency.
- Drive reusable platform services and developer workflows that increase engineering autonomy and reduce operational complexity.
- Provide cross-team technical leadership through architecture influence, mentoring, engineering standards, and hands-on contributions.
- Participate in the on-call rotation and improve platform reliability, observability, and operational excellence.
Requirements
- Experience designing and building large-scale SaaS or cloud platforms with deep distributed systems expertise.
- Strong software architecture and system design skills, with experience leading complex technical initiatives across multiple teams.
- A platform engineering mindset and interest in building reusable capabilities for other engineering organizations.
- Ability to influence technical direction without direct authority and collaborate across organizational boundaries.
- Experience with Java, Go, Kafka, and Kubernetes.
- Experience with observability platforms, internal developer platforms, operational tooling, developer productivity solutions, or AI-powered engineering workflows is beneficial but not required.
Benefits
- Hybrid workplace arrangement
- New hire stock equity (RSUs) and employee stock purchase plan
- Continuous career development and pathing opportunities
- Employee-focused onboarding
- Internal mentor and cross-departmental buddy program
- Friendly and inclusive workplace culture
Tech Stack
Categories
BackendData Engineering
About Datadog
Datadog is the essential monitoring platform for cloud applications. We bring together data from servers, containers, databases, and third-party services to make your stack entirely observable. These capabilities help DevOps teams avoid downtime, resolve performance issues, and ensure customers are getting the best user experience.