OutSystems

Senior Site Reliability Engineer

OutSystems
Apply
3 months ago

Responsibilities

  • Lead and onboard services and teams to reliability tenets.
  • Establish and maintain Service Level Objectives and Service Level Agreements.
  • Design and implement scalable, reliable, secure, and cloud-native infrastructure.
  • Collaborate with software development teams to build resilient, observable, fault-tolerant, recoverable, scalable, and performant systems.
  • Implement monitoring, alerting, logging, and tracing solutions.
  • Lead incident response, resolution, and RCA/post-mortems.
  • Automate operational tasks with emphasis on rapid incident detection and recovery.
  • Develop mission-critical automation and tools in Python using Gen AI tooling.
  • Communicate system reliability and performance updates to stakeholders.
  • Participate in an on-call rotation providing 24/7 production support.

Requirements

  • Bachelor’s or master’s degree in Computer Science or equivalent.
  • At least 6 years of experience in Site Reliability Engineering managing infrastructure and services at scale.
  • History of end-to-end project delivery.
  • Experience managing Hadoop and Kubernetes infrastructure and related services, or equivalent experience.
  • Advanced knowledge of Linux, networking, and containers.
  • Proficiency in at least one high-level programming language such as Python or GoLang.
  • Strong troubleshooting and debugging skills.
  • Fluency in English and excellent written and oral communication skills.
  • Understanding or hands-on experience with prompt engineering in software development.
  • Familiarity with AI-native IDEs or AI assistants such as Cursor, GitHub Copilot, and Claude.
  • Experience with SLOs, SLIs, and SLAs is valued.
  • Experience with Kubernetes, EKS, AWS CloudFormation, Terraform, Puppet, Chef, or Spacelift is valued.
  • Experience with Python, Go, Bash/Shell scripting, or other automation tools and languages is valued.
  • Familiarity with AWS services including EC2, RDS, ELB, CloudFront, and Lambda is valued.
  • Experience with Grafana, ELK, Prometheus, or other monitoring tools is valued.
  • Understanding of resilient and fault-tolerant system design and complex distributed-system troubleshooting is valued; CKA, CKAD, and CKS certifications are valued.

Benefits

  • Hybrid onsite work in Menlo Park, CA.
  • Professional Development Fund and Internal Mobility Program supporting vertical progression, lateral moves, and specialized AI skills.
  • Inclusive, global culture with access to experienced mentors and world-class colleagues.
  • Equal opportunity employment and consideration regardless of protected status.

Tech Stack

Apache HadoopAWSBashChefGoGrafanaKubernetesLinuxPrometheusPuppetPythonTerraform

Categories

Site Reliability
OutSystems

About OutSystems

1,001-5,000 employees
Contact me