
Staff Software Engineer-AI
DigitalOcean3 months ago
Hyderābād, IndiaStaff+
Responsibilities
- Lead the end-to-end design and implementation of multi-agent, multi-turn simulation environments with synthetic users, AI agents, tool integration, and complex workflows.
- Design ML pipelines that analyze historical conversations to generate realistic personas, scenarios, and simulation goals.
- Build evaluation and scoring infrastructure for what-if configuration scenarios and non-deterministic agent outputs.
- Architect high-throughput, stateful workflow orchestration systems for large-scale AI-agent simulations.
- Define scalable API contracts and system boundaries across telemetry data, asynchronous simulation engines, and secure remote execution environments.
- Drive the Feedback Systems technical roadmap and engineering strategy across the Agentic AI organization.
- Solve distributed-systems challenges involving rate limiting, backpressure, state management, and reliable execution of non-deterministic workflows.
- Mentor senior engineers, lead architectural reviews, and establish practices for code quality, testing, and observability.
- Integrate AI/ML platforms, LLMs, prompt routing, and agentic workflows into evaluation infrastructure.
- Partner with product managers and engineering leaders to translate product and experimentation needs into a scalable architecture.
Requirements
- At least 5 years of software engineering experience with modern AI/ML frameworks, LLM orchestration, and production-grade Python and Go.
- At least 10 years of software engineering experience and a proven record operating at Staff, Principal, or Architect level on mission-critical distributed systems.
- Expertise designing highly concurrent, fault-tolerant, and globally scalable backend architectures.
- Deep experience with stateful, durable workflow orchestration engines and complex asynchronous lifecycles at scale.
- Experience designing resilient, high-performance APIs such as gRPC and managing high-throughput message or event-driven architectures.
- Experience processing natural-language data to extract user intent, synthesize personas, and generate deterministic simulation goals.
- Experience building evaluation frameworks for non-deterministic AI systems, including metrics, guardrails, scoring rubrics, and regression testing.
- Experience managing asynchronous state, streaming LLM tokens, handling rate limits, and operating heavy data pipelines.
- Hands-on experience building, scaling, or integrating backend infrastructure for AI/ML products, including LLMs and agentic architectures.
- Strong technical ownership, practical product judgment, and communication skills for collaboration across a globally distributed team.
Benefits
- Hybrid role located in Hyderabad, India.
- Reimbursement for relevant conferences, training, and education.
- Access to LinkedIn Learning's 10,000+ courses.
- Employee Assistance Program, local employee meetups, and flexible time off policy.
- Eligible employees may receive equity grants upon hire and participate in the Employee Stock Purchase Program.
- Bonus and benefits may be available based on company and individual performance and local regulations.
Categories
About DigitalOcean
DigitalOcean provides cloud infrastructure and platform services for developers, startups, and small to mid-sized businesses, including virtual machines (Droplets), managed Kubernetes and databases, object/block storage, networking, and GPUs for AI workloads. It operates a usage-based, self-service public cloud with APIs, CLI, and a marketplace to deploy and scale applications. Founded in 2012 and headquartered in Broomfield, Colorado, DigitalOcean is a public company listed on the NYSE.