
Principal Software Engineer - Observability
The Walt Disney Company25 days ago
Glendale, CA, USA or New York, NY, USAStaff+
Base Salary
$184k - $259k/yr
Responsibilities
- Design and operate intelligent production systems using real-time signals and AI-driven detection to improve streaming platform health, critical services, and customer experience.
- Build and scale autonomous agentic systems using frontier models to detect issues, analyze root causes, and drive resolution with minimal human intervention.
- Develop intelligent harnesses, memory, prompting and context strategies, task decomposition, retrieval, tool routing, multi-agent orchestration, and error recovery patterns.
- Build end-to-end front-end and back-end systems, data and decisioning pipelines, and scalable APIs delivering predictive signals, explainability, and insights.
- Optimize AI usage costs and performance through prompt and context optimization, caching, batching, and smart model routing across hosted and self-hosted models.
- Write clean, tested production code; conduct code reviews; own production components; and work across unfamiliar codebases.
- Embed AI into incident response, release validation, customer insights, observability, reliability, and developer productivity workflows.
- Set technical direction, mentor engineers, and deliver high-impact work tied to business outcomes.
Requirements
- Bachelor’s degree in computer science, engineering, or equivalent experience.
- 10+ years of software engineering experience building AI-powered or data-driven applications, scalable APIs, and production systems at scale.
- Demonstrated expertise in model prompting, context engineering, harnessing, orchestration patterns, and autonomous agentic workflows.
- Strong AI/ML engineering experience orchestrating foundation models such as Claude, OpenAI, and Qwen with frameworks such as LangChain or LangGraph.
- Experience building and deploying systems across both front end and back end on AWS, Azure, or GCP and integrating frontier model APIs.
- Experience with AI-assisted development tools such as Cursor and Claude Code, including prompt optimization, caching, and smart model routing.
- Experience with GitHub, Docker, AWS/EKS, cloud-native deployments, and mature CI/CD pipelines.
- Strong understanding of API design, microservices architecture, and standard SDLC workflows.
- Ability to work independently, set technical direction, mentor engineers, troubleshoot complex issues, and communicate with technical and non-technical stakeholders.
- Preferred experience building observability SaaS solutions and handling high-volume telemetry data with tools such as Datadog, Grafana, Conviva, or OpenTelemetry.
- Preferred experience with AWS Bedrock, local or private deployment of open-source models, inference optimization, and cost/performance tuning.
- Preferred familiarity with PySpark, Pandas, Databricks, Snowflake, prompt design, model evaluation, and foundation-model fine-tuning.
Benefits
- Full-time employment with a stated hiring range of $184,300–$247,100 per year in Glendale, CA or $193,100–$258,900 per year in New York, NY.
- Additional bonus and/or long-term incentive units may be provided.
- Medical, financial, and other benefits may be included depending on the level and position offered.
- The role is based in New York, NY, with an alternate location in Glendale, CA.