22 days ago
São Paulo, BrazilSenior
Responsibilities
- Continuously improve the reliability and performance of Latitude.sh’s bare metal cloud platform.
- Design, build, and maintain tools that automate operational tasks and incident response.
- Implement and improve monitoring, alerting, tracing, and other observability solutions.
- Collaborate with engineering and platform teams on scalable and resilient systems.
- Participate in on-call rotations and lead learning-focused post-incident reviews.
- Develop and document operational processes and runbooks.
- Contribute to SLO and SLI definitions and adoption of reliability metrics.
Requirements
- Strong verbal and written English communication skills.
- Advanced knowledge of Linux/Unix systems in production environments.
- Experience with Kubernetes and container orchestration.
- Proficiency with infrastructure automation tools such as Terraform and Ansible.
- Experience with observability stacks such as Prometheus, Grafana, Loki, and ELK.
- Familiarity with Bash, Python, Go, or Ruby.
- Working knowledge of Git and CI/CD pipelines.
- Understanding of incident management and root cause analysis processes.
- Knowledge of cloud-native reliability and security best practices.
Benefits
- Contractor (PJ).
- Paid time off.
- Competitive compensation.
- Wellhub, formerly Gympass.
- Annual bonus based on company and team performance.
- Flexible work hours.
- Professional growth and development opportunities.
Categories
Site Reliability
