
Senior Site Reliability Engineer
Salesforce2 days ago
Dublin, IrelandSenior
Responsibilities
- Lead detection, response, resolution, root cause analysis, postmortems, and corrective actions for production incidents.
- Design and implement automation platforms, self-healing systems, durable workflow pipelines, and AI-powered operational tooling.
- Architect and build production observability solutions covering monitoring, logging, alerting, tracing, and autonomous remediation.
- Develop AI/ML-powered operations tools such as anomaly detection, predictive analysis pipelines, intelligent runbook automation, and prompt-engineered operational agents.
- Optimize system performance, reliability, availability, and cost through proactive monitoring and tuning.
- Define and uphold SLAs and SLOs, use error budgets, reduce toil, and improve time to detect and restore services.
- Collaborate with product and engineering teams to design, build, and operate reliable systems.
- Ensure Site Reliability work complies with Salesforce internal compliance policies and directives.
- Create technical epics with clear problem statements, implementation documentation, and measurable business outcomes.
- Build and ship production-grade software, evaluate human- and AI-generated code, and maintain shared system context for reliable AI-assisted engineering.
- Coach junior engineers through pair programming, design reviews, and code reviews.
- Work in a 24/7 global operations model, including weekend on-call coverage.
Requirements
- 5+ years of systems engineering and software engineering experience supporting large-scale, internet-facing services.
- Hands-on experience with Docker, Kubernetes, and containerized orchestration platforms.
- Strong knowledge of distributed systems, Linux/Unix internals, performance tuning, and large-scale troubleshooting.
- Familiarity with DNS, HTTP, load balancing, caching, and large-scale internet service architectures.
- Proficiency in Python and Go with strong software engineering practices.
- Production experience with observability platforms such as Grafana, Prometheus, ELK, Splunk, or Datadog.
- Experience with incident management, on-call operations, root cause analysis, and postmortems.
- Understanding of SRE principles including SLIs, SLOs, error budgets, toil reduction, blameless culture, and capacity planning.
- Hands-on experience with Temporal, Airflow, Argo Workflows, or similar workflow and orchestration engines.
- Experience applying AI/ML to operations, including anomaly detection, predictive analysis, LLM-based automation, and prompt engineering.
- Experience using AI development tools such as Claude Code, GitHub Copilot, Codex, or Cursor.
- Advanced prompt engineering skills and ability to maintain reliable, secure, production-ready system context.
- Excellent communication, incident leadership, technical presentation, mentoring, and technical coaching skills.
- Ability to manage multiple priorities in a time-sensitive 24/7 global operations environment.
- A related technical degree is required.
- Preferred qualifications include AI agent frameworks, MCP, LLM-powered operational tools, open-source reliability or observability contributions, AWS or GCP professional-level certifications, multi-cloud or hyperscale SRE experience, and chaos engineering or game day experience.
Benefits
- Salesforce offers benefits and resources intended to support balance and employee well-being.
- The role is based in Dublin and operates five days a week, 24 hours a day, in a follow-the-sun model with weekend on-call coverage.
Tech Stack
Categories
AI ApplicationsSite Reliability
About Salesforce
Salesforce builds cloud-based customer relationship management software and a broader Customer 360 platform for sales, service, marketing, commerce, and analytics, sold by subscription to businesses and public-sector organizations. Founded in 1999 and headquartered in San Francisco, it trades on the NYSE under the symbol CRM. Its portfolio includes Sales Cloud, Service Cloud, Marketing Cloud, MuleSoft integration, Tableau analytics, and Slack for collaboration, with extensive developer tools and APIs.