4 months ago
Remote, BrazilSenior
Responsibilities
- Embed with critical product areas to assess architecture, reliability risks, and operational maturity.
- Implement SRE practices including SLOs, on-call structures, runbooks, incident management, and post-mortems.
- Build and evolve SRE maturity matrices and identify risks such as single points of failure, ownerless services, and noisy alerts.
- Promote adoption of the Agentic Engineering Platform, Golden Paths, canary deployments, feature flags, and automated dashboards.
- Develop monitoring, observability, and alerting solutions and support SLIs, SLOs, and error budgets.
- Partner with Engineering Platform and product teams to improve platform adoption and reliability culture.
- Use data on availability, latency, toil, and infrastructure usage to guide decisions.
- Support critical-system migrations, environment segregation, and deprecation of legacy technologies.
Requirements
- Experience with cloud environments, preferably GCP.
- Proficiency with observability tools and practices, including Prometheus, Grafana, Loki, Thanos, Elasticsearch, and AlertManager.
- Strong knowledge of Kubernetes and distributed-systems architecture.
- Strong knowledge of infrastructure as code and Terraform.
- Hands-on experience with incident management, on-call operations, and post-mortems.
- Experience defining and tracking SLOs and error budgets.
- Ability to analyze logs and distributed-system performance.
- Strong communication and influence skills across engineers, product managers, and leadership.
- Data-driven approach to mapping risks, prioritizing actions, and demonstrating impact.
- Preferred: multi-cloud or high-traffic company experience, software development experience, embedded or product-team SRE experience, observability/toil-reduction/operations-automation contributions, engineering-platform concepts, or incident.io and Grafana IRM experience.
Benefits
- Fully remote work model with freedom to choose where to live.
- Work with an internal Agentic Engineering Platform and large-scale legal-information systems.
Tech Stack
Categories
Site Reliability
