Responsibilities
- Design and build production AI features such as LLM agents, retrieval systems, and event intelligence for high-volume real-time event streams.
- Architect agent and prompt orchestration, retrieval pipelines, tool and API integrations, low-latency inference, and evaluation systems.
- Own the AI feature lifecycle from problem framing and prototyping through production deployment, monitoring, guardrails, and improvement loops.
- Design for consistency, throughput, fault tolerance, reliability, and cost under bursty and unpredictable loads.
- Partner with platform, product, and applied-research teams to define success criteria and integrate AI into existing services.
- Provide technical mentorship, participate in reviews, and help shape the team’s technical direction.
Requirements
- At least 5 years of software engineering experience, including meaningful experience building and operating production distributed systems.
- Hands-on experience building and shipping production AI systems, including LLM-powered applications, agents, or retrieval systems and their orchestration, serving, and evaluation components.
- Strong programming fundamentals and the ability to work across systems and AI/application code.
- Applied AI knowledge covering prompting, retrieval, agent patterns, LLM evaluation, and behavioral guardrails.
- Experience with cloud infrastructure such as AWS, GCP, or Azure, as well as containers and Kubernetes orchestration.
- Reliability-minded systems judgment and the ability to explain technical trade-offs.
- Strong communication and collaboration skills with a record of improving teams and systems.
- Preferred: experience with LLMOps, evaluation harnesses, prompt and version management, agent tracing and observability, and online/offline evaluation consistency.
- Preferred: experience with production LLM serving, retrieval-augmented generation, multi-step agents, tool use, anomaly detection, event correlation, observability, AIOps, or reliability.
- Preferred: familiarity with LLM APIs, LangChain, LlamaIndex, vector databases, Kafka, Airflow, or Spark, plus open-source AI or distributed-systems contributions.
Benefits
- Hybrid work model with flexibility within PagerDuty’s established office locations in Atlanta, Lisbon, London, San Francisco, Santiago, Sydney, Tokyo, and Toronto; candidates must reside in an eligible location.
- Competitive salary and comprehensive benefits package.
- Flexible work arrangements, company equity, and an Employee Stock Purchase Program, subject to eligibility.
- Retirement or pension plan, generous paid vacation, paid holidays, and sick leave.
- Dutonian Wellness Days and HibernationDuty companywide paid days off in addition to PTO.
- Paid parental leave of 22 weeks for a pregnant parent and 12 weeks for a non-pregnant parent, subject to local laws and eligibility.
- 20 hours of paid volunteer time off per year.
- Company-wide hack weeks and mental wellness programs.
Tech Stack
Categories
About PagerDuty
In an always-on world, teams trust PagerDuty to help them deliver an optimal digital experience to their customers, every time. PagerDuty is the central nervous system for a company’s digital operations. We identify issues and opportunities in real-time and bring together the right people to respond to problems faster and prevent them in the future. From digital disruptors to nearly half of the Fortune 500 and the Forbes AI 50s, over 34,000 paid and free customers rely on PagerDuty to help them continually improve their digital operations—so their teams can spend less time reacting to incidents and more time building for the future. Dutonians believe that we are a part of a bigger movement of businesses being built to benefit everyone—the customer and the employee, as well as our community. We are go-getters fueled by the fire to reinvent how people and companies work together. We take the lead and get creative to be first in the hearts of our customers. Whether it’s keeping the world on or changing it entirely, Dutonians are fueled by the fire to reinvent how people and companies work together to deliver in real-time, across the globe. Join us to lead uncharted efforts and reinvent how companies run.