1 day ago
London, United KingdomMid Level / Senior
Responsibilities
- Design and implement resilient software systems for AI-assisted observability, automated incident response, and self-healing cloud infrastructure.
- Build and maintain production-grade AI agents, Model Context Protocol integrations, and deterministic evaluation pipelines.
- Develop ingestion and correlation pipelines for logs, metrics, traces, change events, and runbooks.
- Create anomaly detection and human-in-the-loop remediation workflows with safety, security, and quality guardrails.
- Define SLIs and SLOs, manage error budgets, and lead post-incident reviews.
- Mentor engineers, conduct code reviews, establish engineering guidelines, and drive operational excellence across teams.
- Coordinate priorities and delivery across application, infrastructure, development, and operations teams.
Requirements
- Bachelor’s degree and 8 years of related experience, master’s degree and 6 years, or PhD and 3 years in computer science, software engineering, or a related technical field.
- Experience as a Senior or Lead SRE or Software Engineer delivering distributed, highly available SaaS platforms at scale.
- Strong proficiency in Python, Go, Java, or C++ and experience designing production automation and distributed services.
- Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments.
- Experience with SRE practices, SLI/SLO design, observability platforms, incident management, and automated root-cause analysis.
- Preferred experience building LLM pipelines, AI agents, Model Context Protocol servers or clients, RAG architectures, and evaluation frameworks.
- Preferred experience with OpenTelemetry, Prometheus, Grafana, Splunk, ThousandEyes, or distributed tracing systems.
- Preferred expertise with AWS, GCP, Azure, Terraform, Jenkins, GitHub Actions, or Harness.
- Preferred experience with responsible AI guardrails, deterministic fallback logic, policy-driven remediation, Kafka, Redis, PostgreSQL, and Elasticsearch or vector databases.
Tech Stack
Apache KafkaAWSAzureC++DockerElasticsearchGitHub ActionsGoGoogle Cloud PlatformGrafanaHarnessJavaJenkinsKubernetesPostgreSQLPrometheusPythonRedisSplunkTerraform
Categories
AI ApplicationsSite Reliability
About Cisco
Cisco designs and sells networking, security, and collaboration platforms for enterprises, service providers, and governments, spanning routers and switches, Wi‑Fi, firewalls, zero‑trust, observability, and cloud-managed IT (Meraki) plus Webex. Its business model mixes hardware, software subscriptions, and support/consulting services. Founded in 1984 and headquartered in San Jose, California, Cisco is a public company traded on Nasdaq and serves customers across data centers, campuses, and service provider networks.
