1 day ago
Base Salary
$168k - $245k/yr
Responsibilities
- Define and enforce feature-level SLIs, SLOs, error budgets, and burn-rate alerting for APIs, RAG systems, AI agents, and user-facing applications.
- Build application observability systems and Looker dashboards using BigQuery and BigTable to track feature health, errors, and usage trends.
- Design LangGraph-based agents for anomaly detection, root-cause diagnosis, automated rollback, feature-flag kill switches, and self-healing workflows.
- Develop evaluation harnesses for agent benchmarking, multi-step workflows, non-deterministic outputs, and regression testing.
- Write complex BigQuery SQL and design schemas for usage analysis, anomaly detection, operational analytics, observability, and debugging.
- Analyze adoption and usage trends to identify reliability risks, capacity needs, and degraded user experiences.
- Partner with application teams on deployment safety, structured logging, distributed tracing, and reliability practices.
- Lead application-level incident response, root-cause analysis, and blameless postmortems.
- Build Python tooling and automation to reduce mean time to detect and mean time to resolve application issues.
- Apply emerging AI frameworks, tools, and techniques to improve platform reliability and developer productivity.
Requirements
- 7+ years of software engineering experience with significant focus on reliability, observability, or production operations.
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical discipline.
- Strong Python development skills, including production tooling, automation, and agent-based systems.
- Production GCP experience deploying and managing applications on GKE and Kubernetes.
- Deep BigQuery SQL expertise, including complex queries, window functions, schema design, and cost optimization.
- Hands-on BigTable or equivalent high-throughput operational data experience.
- Experience designing and operating application-level SLI/SLO frameworks, burn-rate alerting, and error budget policies.
- Strong application-layer debugging skills involving distributed tracing, profiling, structured log analysis, and dependency mapping.
- Preferred experience with agent evaluation harnesses, A2A protocols, streaming architectures, event-driven systems, feature flags, canary deployments, progressive rollouts, and automated rollback.
- Preferred experience with Cloud Logging, Cloud Trace, Cloud Monitoring, AIOps, ML-driven anomaly detection, automated root-cause analysis, intelligent alerting, and reliability programs.
- Preferred hands-on GenAI application development experience with LangGraph, agent engineering, prompt design, and agentic workflows.
- Preferred experience building Looker dashboards and LookML models for operational observability.
Benefits
- Hybrid work model based in San Jose, California or North Carolina.
- Starting salary range of $167,700.00 to $245,200.00, excluding incentive compensation, equity, and benefits.
- Medical, dental, and vision insurance; 401(k) with Cisco matching contribution; paid parental leave; disability coverage; and basic life insurance.
- Potential eligibility for Cisco restricted stock units and annual bonuses for non-sales roles.
- Paid holidays, floating holiday, birthday day off, year-end shutdown, personal wellness days, vacation or flexible vacation time, sick time, family emergency leave, and optional paid volunteer days.
Tech Stack
Categories
Site Reliability
About Cisco
Cisco designs and sells networking, security, and collaboration platforms for enterprises, service providers, and governments, spanning routers and switches, Wi‑Fi, firewalls, zero‑trust, observability, and cloud-managed IT (Meraki) plus Webex. Its business model mixes hardware, software subscriptions, and support/consulting services. Founded in 1984 and headquartered in San Jose, California, Cisco is a public company traded on Nasdaq and serves customers across data centers, campuses, and service provider networks.
