
Site Reliability Engineer
CSC Generation2 months ago
San José, Costa RicaMid Level
Responsibilities
- Improve service resiliency, performance, scalability, and system design across Backcountry’s platform.
- Drive critical incident resolution and implement fixes through structured postmortems.
- Design and implement automation to reduce operational toil.
- Build and maintain observability using metrics, logs, traces, profiles, and SLI/SLO instrumentation.
- Monitor system health and capacity and take proactive corrective action.
- Collaborate with developers, architects, DevOps, and IT operations teams to build, deploy, and support features.
- Participate in FinOps initiatives across GCP and AWS, including capacity planning and committed-use discount strategy.
- Participate in the SRE on-call support rotation.
- Use AI-assisted engineering tools to investigate issues, automate work, perform code reviews, and ship infrastructure and application fixes.
Requirements
- At least 3 years of experience supporting containerized production services, preferably with Kubernetes.
- At least 3 years of experience with infrastructure as code such as Terraform, AWS CDK, or Ansible.
- At least 3 years of cloud experience operating in Google Cloud Platform and/or AWS.
- Ability to diagnose issues and ship fixes directly to application code and investigate infrastructure and application repositories end to end.
- Proficiency with Bash, Python, and TypeScript/Node.js.
- Production Linux administration experience.
- Knowledge of DHCP, DNS, HTTPS, SSH, DevOps, and SRE practices.
- Hands-on experience with observability tools such as Grafana, Prometheus, Loki, or OpenSearch, plus SLI/SLO instrumentation.
- Experience with GitOps and Kubernetes packaging using ArgoCD, Helm, or Kustomize.
- Bachelor’s degree in computer science or a similar field, or equivalent experience.
- Advanced-level English communication skills.
- Preferred qualifications include AI coding assistant experience, regulated or PCI-scoped environment experience, ecommerce experience, and GCP, CKA, or AWS certifications.
Benefits
- Primarily remote work in Costa Rica, with a mandatory in-person interview before an offer.
- Private medical and life insurance.
- Additional paid time off, monthly allowances, and reimbursements.
- Employee discounts and outdoor-related employee perks.
- Opportunities for professional growth.
- Cross-functional ownership and direct impact on platform reliability.
Tech Stack
AnsibleAWSAzureBashGoogle Cloud PlatformGrafanaHelmKubernetesLinuxNode.jsPrometheusPythonTerraformTypeScript
Categories
Site Reliability