
Senior Site Reliability Engineer
OutSystems3 months ago
Responsibilities
- Lead and onboard services and teams to reliability tenets.
- Establish and maintain Service Level Objectives and Service Level Agreements.
- Design and implement scalable, reliable, secure, and cloud-native infrastructure.
- Collaborate with software development teams to build resilient, observable, fault-tolerant, recoverable, scalable, and performant systems.
- Implement monitoring, alerting, logging, and tracing solutions.
- Lead incident response, resolution, and RCA/post-mortems.
- Automate operational tasks with emphasis on rapid incident detection and recovery.
- Develop mission-critical automation and tools in Python using Gen AI tooling.
- Communicate system reliability and performance updates to stakeholders.
- Participate in an on-call rotation providing 24/7 production support.
Requirements
- Bachelor’s or master’s degree in Computer Science or equivalent.
- At least 6 years of experience in Site Reliability Engineering managing infrastructure and services at scale.
- History of end-to-end project delivery.
- Experience managing Hadoop and Kubernetes infrastructure and related services, or equivalent experience.
- Advanced knowledge of Linux, networking, and containers.
- Proficiency in at least one high-level programming language such as Python or GoLang.
- Strong troubleshooting and debugging skills.
- Fluency in English and excellent written and oral communication skills.
- Understanding or hands-on experience with prompt engineering in software development.
- Familiarity with AI-native IDEs or AI assistants such as Cursor, GitHub Copilot, and Claude.
- Experience with SLOs, SLIs, and SLAs is valued.
- Experience with Kubernetes, EKS, AWS CloudFormation, Terraform, Puppet, Chef, or Spacelift is valued.
- Experience with Python, Go, Bash/Shell scripting, or other automation tools and languages is valued.
- Familiarity with AWS services including EC2, RDS, ELB, CloudFront, and Lambda is valued.
- Experience with Grafana, ELK, Prometheus, or other monitoring tools is valued.
- Understanding of resilient and fault-tolerant system design and complex distributed-system troubleshooting is valued; CKA, CKAD, and CKS certifications are valued.
Benefits
- Hybrid onsite work in Menlo Park, CA.
- Professional Development Fund and Internal Mobility Program supporting vertical progression, lateral moves, and specialized AI skills.
- Inclusive, global culture with access to experienced mentors and world-class colleagues.
- Equal opportunity employment and consideration regardless of protected status.
Categories
Site Reliability