Informatica

Senior Site Reliability Engineer

Informatica
Apply
2 months ago

Base Salary

$149k - $224k/yr

Responsibilities

  • Lead incident detection, response, resolution, root cause analysis, postmortems, and corrective actions for customer-facing services.
  • Design and implement automation platforms, self-healing systems, durable workflow pipelines, and AI-powered operational tooling.
  • Build production-grade observability systems covering monitoring, logging, alerting, tracing, proactive detection, and autonomous remediation.
  • Develop AI/ML-powered operations tools, including anomaly detection, predictive analysis, intelligent runbook automation, and MCP-based operational agents.
  • Improve system performance, reliability, availability, and cost-effectiveness while reducing toil and operational overhead.
  • Define and uphold SLAs and SLOs, use error budgets to guide decisions, and drive service reliability improvements.
  • Collaborate with product and engineering teams to design, build, and operate reliable systems from the beginning of the lifecycle.
  • Lead technical epics and produce problem statements, implementation documentation, and measurable business outcomes.
  • Mentor junior engineers through pair programming, design reviews, code reviews, and technical coaching.
  • Build and ship secure, optimized, production-grade software using AI-assisted development tools and evaluate human- and AI-generated code for correctness, quality, security, and performance.
  • Maintain shared system context containing system designs, constraints, and standards for reliable AI-assisted engineering.
  • Work within a 24/7 global operations model, including weekend on-call responsibilities.

Requirements

  • At least 5 years of experience in systems engineering and software engineering for large-scale, internet-facing services.
  • Required related technical degree.
  • Hands-on experience with Docker, Kubernetes, and containerized orchestration platforms.
  • Strong knowledge of distributed systems, Linux/Unix internals, and large-scale performance troubleshooting.
  • Familiarity with DNS, HTTP, load balancing, caching, and large-scale internet service architectures.
  • Proficiency in Python and Go with strong software engineering practices, including testing, code review, and CI/CD.
  • Production experience with observability platforms such as Grafana, Prometheus, ELK, Splunk, or Datadog.
  • Experience with incident management, on-call operations, root cause analysis, and postmortems.
  • Strong understanding of SLIs, SLOs, error budgets, toil reduction, blameless culture, and capacity planning.
  • Hands-on experience with Temporal, Airflow, Argo Workflows, or similar workflow and orchestration engines.
  • Experience applying AI/ML to operations, including anomaly detection, predictive analysis, LLM-based automation, and prompt engineering.
  • Excellent communication, incident leadership, technical presentation, mentoring, and prioritization skills.
  • Ability to work in a 24/7 global operations environment and manage time-sensitive priorities.
  • Experience using AI development tools such as Claude Code, GitHub Copilot, Codex, or Cursor.
  • Advanced prompt engineering skills and ability to develop reliable, secure, production-ready AI workflows.
  • Preferred: experience with AI agent frameworks, MCP, LLM-powered operational tools, open-source reliability or observability tooling, chaos engineering, game day exercises, multi-cloud or hyperscale SRE environments, and AWS or GCP professional-level certifications.

Benefits

  • Benefits include time off programs, medical, dental, vision, mental health support, paid parental leave, life and disability insurance, 401(k), and an employee stock purchasing program.
  • The role operates in a 24/7 global follow-the-sun model with weekend on-call responsibilities.
  • Salesforce provides reasonable accommodation support during the application and recruiting process.

Tech Stack

Apache AirflowAWSDatadogDockerGoGoogle Cloud PlatformGrafanaKubernetesLinuxPrometheusPythonSplunk

Categories

Site Reliability
Informatica

About Informatica

5,001-10,000 employees

Informatica (NYSE: INFA), a leader in AI-powered enterprise cloud data management, helps businesses unlock the full value of their data and AI. As data grows in complexity and volume, Informatica’s Intelligent Data Management Cloud™ delivers a complete, end-to-end platform with a suite of industry-leading, integrated solutions to connect, manage and unify data across any cloud, hybrid or multi-cloud environment. Powered by CLAIRE® AI, Informatica’s platform integrates natively with all major cloud providers, data warehouses and analytics tools— giving organizations the freedom of choice, avoiding vendor lock-in and delivering better ROI by enabling access to governed data, simplifying operations and scaling with confidence. Trusted by approximately 5,000 customers in nearly 100 countries—including over 80 of the Fortune 100—Informatica is the backbone of platform-agnostic, cloud data-driven transformation. Informatica. Where data and AI come to life.™ Contact: PR@informatica.com

Contact me