DigitalOcean

Staff Software Engineer-AI

DigitalOcean
Apply
3 months ago
Hyderābād, IndiaStaff+

Responsibilities

  • Lead the end-to-end design and implementation of multi-agent, multi-turn simulation environments with synthetic users, AI agents, tool integration, and complex workflows.
  • Design ML pipelines that analyze historical conversations to generate realistic personas, scenarios, and simulation goals.
  • Build evaluation and scoring infrastructure for what-if configuration scenarios and non-deterministic agent outputs.
  • Architect high-throughput, stateful workflow orchestration systems for large-scale AI-agent simulations.
  • Define scalable API contracts and system boundaries across telemetry data, asynchronous simulation engines, and secure remote execution environments.
  • Drive the Feedback Systems technical roadmap and engineering strategy across the Agentic AI organization.
  • Solve distributed-systems challenges involving rate limiting, backpressure, state management, and reliable execution of non-deterministic workflows.
  • Mentor senior engineers, lead architectural reviews, and establish practices for code quality, testing, and observability.
  • Integrate AI/ML platforms, LLMs, prompt routing, and agentic workflows into evaluation infrastructure.
  • Partner with product managers and engineering leaders to translate product and experimentation needs into a scalable architecture.

Requirements

  • At least 5 years of software engineering experience with modern AI/ML frameworks, LLM orchestration, and production-grade Python and Go.
  • At least 10 years of software engineering experience and a proven record operating at Staff, Principal, or Architect level on mission-critical distributed systems.
  • Expertise designing highly concurrent, fault-tolerant, and globally scalable backend architectures.
  • Deep experience with stateful, durable workflow orchestration engines and complex asynchronous lifecycles at scale.
  • Experience designing resilient, high-performance APIs such as gRPC and managing high-throughput message or event-driven architectures.
  • Experience processing natural-language data to extract user intent, synthesize personas, and generate deterministic simulation goals.
  • Experience building evaluation frameworks for non-deterministic AI systems, including metrics, guardrails, scoring rubrics, and regression testing.
  • Experience managing asynchronous state, streaming LLM tokens, handling rate limits, and operating heavy data pipelines.
  • Hands-on experience building, scaling, or integrating backend infrastructure for AI/ML products, including LLMs and agentic architectures.
  • Strong technical ownership, practical product judgment, and communication skills for collaboration across a globally distributed team.

Benefits

  • Hybrid role located in Hyderabad, India.
  • Reimbursement for relevant conferences, training, and education.
  • Access to LinkedIn Learning's 10,000+ courses.
  • Employee Assistance Program, local employee meetups, and flexible time off policy.
  • Eligible employees may receive equity grants upon hire and participate in the Employee Stock Purchase Program.
  • Bonus and benefits may be available based on company and individual performance and local regulations.

Tech Stack

DigitalOcean

About DigitalOcean

1,001-5,000 employees

DigitalOcean provides cloud infrastructure and platform services for developers, startups, and small to mid-sized businesses, including virtual machines (Droplets), managed Kubernetes and databases, object/block storage, networking, and GPUs for AI workloads. It operates a usage-based, self-service public cloud with APIs, CLI, and a marketplace to deploy and scale applications. Founded in 2012 and headquartered in Broomfield, Colorado, DigitalOcean is a public company listed on the NYSE.

Contact me