3 months ago
Base Salary
$170k - $276k/yr
Responsibilities
- Design and operate the agent execution runtime for scheduling, state management, and lifecycle management of long-running workflows
- Build multi-agent coordination capabilities covering task handoff, agent memory, tool use, and workflow branching
- Develop context-window management and session-persistence layers for stateful agent interactions
- Build prompt engineering tools for prompt versioning, testing, and evaluation at scale
- Support AI-assisted coding workflows through IDE integrations, local-first development environments, and fast iteration loops
- Own Agentic Platform services, including AWS or OCI provisioning, scaling, cost management, and reliability
- Provision infrastructure with Terraform and package and deploy Kubernetes applications with Helm
- Participate in a 24/7 on-call rotation and take pager responsibility for owned services
- Define SLOs, lead incident response and blameless post-mortems, and drive reliability improvements
- Implement observability through distributed tracing, structured logging, metrics dashboards, and alerting
- Set technical direction, make architectural decisions, mentor engineers, and collaborate with product, design, and engineering partners
Requirements
- 8+ years of professional, hands-on, full-time software engineering experience in backend, infrastructure, or platform engineering
- Hands-on experience operating production services in AWS, Oracle Cloud Infrastructure, Azure, or GCP, including compute, networking, managed services, IAM, and cost management
- End-to-end production service ownership, including on-call, incident response, SLO definition, and post-mortems
- Deep understanding of fault tolerance, consistency, observability, and scalability in cloud-native distributed systems
- Strong proficiency in at least one backend language used for systems work: Go, Python, Rust, or Java
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience
- Professional proficiency in Go is strongly preferred
- Experience with Terraform and Helm is strongly preferred
- Experience with PostgreSQL and Redis or Pub/Sub patterns is strongly preferred
- Experience with MCP server design and integration is strongly preferred
- Experience with Docker, Kubernetes, or equivalent container and orchestration technologies is strongly preferred
- Familiarity with Cursor, Claude Code, Copilot, Windsurf, or similar AI-assisted development tools is strongly preferred
- Experience with LLM-as-judge frameworks, behavioral regression testing, and golden dataset management is strongly preferred
- Hands-on experience building or operating AI agent systems, including multi-agent orchestration, tool use, memory systems, or agent evaluation frameworks, is strongly preferred
- Open-source contributions or community engagement are strongly preferred
Benefits
- Remote-first work with flexibility and offices in Seattle and Paris
- Quarterly Whaleness Days and an end-of-year Whaleness break
- Home office setup
- 16 weeks of paid parental leave after six months of employment
- Technology stipend of $100 USD net per month
- PTO plan
- Training stipend for conferences, courses, and classes
- Equity
- Medical benefits, retirement benefits, and holidays varying by country
- Docker Swag
Tech Stack
Categories
About Docker
At Docker, we simplify the lives of developers who are making world-changing apps. Docker helps developers bring their ideas to reality by conquering the complexity of app development. We simplify and accelerate workflows with an integrated development pipeline and application components. Actively used by millions of developers around the world, Docker Desktop and Docker Hub provide unmatched simplicity, agility and choice.
