
Staff Site Reliability Engineer, Infrastructure Platform
Procore Technologies2 hours ago
Cairo, EgyptStaff+
Responsibilities
- Lead the design and implementation of the next-generation platform, focusing on distributed systems, internal developer experience, and ease of adoption.
- Build, maintain, and support cloud-based compute platform capabilities for SaaS products.
- Create documentation, lead training sessions, and share best practices to enable internal engineering teams.
- Drive software development lifecycle activities while ensuring platform stability, observability, and compliance.
- Identify platform gaps, run proofs of concept and experiments, and evaluate advancements in the Kubernetes ecosystem.
- Partner with Product, UX, and IT teams to influence the roadmap and improve developer productivity.
- Stay current on platform and Kubernetes trends and advocate for appropriate adoption.
- Master generative tools and agentic workflows and contribute to building agentic capabilities.
Requirements
- 6+ years of experience maintaining application services in Kubernetes clusters.
- Experience provisioning and operating cloud-native tools at scale.
- 3+ years of experience as a CKD.
- 6+ years of experience operating highly available, distributed, scalable cloud-based systems with large amounts of data.
- Experience composing and maintaining Kubernetes manifests, including Helm and Kustomize.
- Experience managing infrastructure as code, including Terraform.
- Experience developing and customizing CI/CD pipeline automation and building container images.
- Experience with AWS, including RDS, EC2, S3, and IAM.
- Ability to lead technical enablement initiatives, mentor engineers, and drive adoption of new platform features.
Tech Stack
Categories
DevOpsSite Reliability