Responsibilities
- Design and implement intelligent automation for infrastructure lifecycle management, including self-healing, anomaly detection, and automated remediation.
- Apply AI/ML techniques for predictive monitoring and proactive performance optimization.
- Lead complex incident response efforts and root cause analyses.
- Improve system reliability through dynamic scaling, telemetry instrumentation, and automated performance tuning.
- Enhance operational runbooks and playbooks by eliminating manual processes through automation.
- Evaluate and adopt emerging AIOps, cloud-native, and distributed systems technologies.
- Partner with Product Engineering, Customer Support, Global Services, Deal Desk, and Sales to deliver strong customer experiences.
Requirements
- 2–4 years of experience in Linux systems administration and/or Python development in production environments.
- Strong Linux administration skills, including troubleshooting, service management, performance tuning, and networking fundamentals.
- Experience developing Python scripts or lightweight applications for operational automation and system management.
- Hands-on Docker experience and familiarity with Kubernetes concepts such as deployments, services, and scaling.
- At least one year of experience supporting SaaS or cloud-native production environments.
- Working knowledge of messaging platforms and databases such as Kafka, Redis, and MySQL or similar technologies.
- Experience contributing to CI/CD pipelines and deployment automation.
- Hands-on experience with monitoring and observability platforms such as Prometheus and Grafana or similar tools.
- Experience participating in incident response, post-incident reviews, and root cause analysis.
- Experience with Jenkins, Terraform, GitOps, or advanced Infrastructure as Code practices is preferred.
- Exposure to AI/ML technologies for anomaly detection, predictive operations, or intelligent automation is preferred.
- Relevant certifications such as RHCSA, AWS, Azure, GCP, PCAP, Docker Certified Associate, Certified Kubernetes Administrator, or SRE-related certifications are preferred.
Benefits
- Hybrid work model in Heredia, Costa Rica, with 3 days in the office and 2 days remote
- Medical, dental, and vision insurance
- Flexible paid time off, holidays, wellness days, and a company-wide year-end break
- Paid parental leave, subject to local policy
- Learning and development stipend
- Volunteer opportunities and charitable donation matching where available
- Mental wellbeing resources and support
Tech Stack
Categories
About Zuora
Zuora was born out of a vision that we could evangelize a fundamentally new way of doing business by shifting the focus of companies to deliver recurring, people-centric services instead of a one-time sale of products. This is how we coined the term, the Subscription Economy®. Today, we see others evangelizing this term, and building entire communities around it. The Subscription Economy isn’t (and never was) just about subscription business models but, direct, recurring relationships with customers through any business model. Subscriptions were only just scratching the surface and now, the market recognizes the Subscription Economy for what it truly is-a relationship-centric economy. Companies have realized that the path to growth going forward is to establish direct, digital relationships with their customers, and to nurture and monetize these relationships through an ever growing set of digital services. Alongside this evolution, Zuora has been there every step of the way. We started with Zuora Billing, and have expanded our award-winning multi-product portfolio to include Zuora Revenue, Zuora Payments and Zuora Platform. More recently, we’ve added Zephr and Togai to our family, further expanding our capabilities to serve as an intelligent hub that monetizes the complete quote to cash and revenue recognition process at scale. We call this Monetization.