Megaport

Senior Site Reliability Engineer

Megaport
Apply
22 days ago
São Paulo, BrazilSenior

Responsibilities

  • Continuously improve the reliability and performance of Latitude.sh’s bare metal cloud platform.
  • Design, build, and maintain tools that automate operational tasks and incident response.
  • Implement and improve monitoring, alerting, tracing, and other observability solutions.
  • Collaborate with engineering and platform teams on scalable and resilient systems.
  • Participate in on-call rotations and lead learning-focused post-incident reviews.
  • Develop and document operational processes and runbooks.
  • Contribute to SLO and SLI definitions and adoption of reliability metrics.

Requirements

  • Strong verbal and written English communication skills.
  • Advanced knowledge of Linux/Unix systems in production environments.
  • Experience with Kubernetes and container orchestration.
  • Proficiency with infrastructure automation tools such as Terraform and Ansible.
  • Experience with observability stacks such as Prometheus, Grafana, Loki, and ELK.
  • Familiarity with Bash, Python, Go, or Ruby.
  • Working knowledge of Git and CI/CD pipelines.
  • Understanding of incident management and root cause analysis processes.
  • Knowledge of cloud-native reliability and security best practices.

Benefits

  • Contractor (PJ).
  • Paid time off.
  • Competitive compensation.
  • Wellhub, formerly Gympass.
  • Annual bonus based on company and team performance.
  • Flexible work hours.
  • Professional growth and development opportunities.

Tech Stack

AnsibleBashGitGoGrafanaKubernetesLinuxPrometheusPythonRubyTerraform

Categories

Site Reliability
Megaport

About Megaport

501-1,000 employees
Contact me