2 months ago
Base Salary
$160k - $200k/yr
Responsibilities
- Design and implement New Relic monitoring, alerting, dashboards, and NRQL queries across Azure and AWS.
- Define SLOs, SLIs, and error budgets and coach teams on reliability practices and system health communication.
- Reduce alert noise, improve signal quality, optimize observability costs, and establish logging, metrics, tracing, and dashboard standards.
- Develop and maintain Terraform for monitoring resources and observability infrastructure and enforce IaC governance.
- Author and troubleshoot Azure DevOps pipelines and improve deployment visibility, change tracking, and release hygiene.
- Administer Incident.io, including alert routing, notification workflows, Slack and OpsGenie integrations, and runbooks.
- Establish incident management foundations including postmortems, on-call rotations, escalation policies, severity classifications, and response playbooks.
- Respond to and debrief production incidents, track MTTR and MTTD, and drive continuous improvement.
- Enable engineering teams through workshops, consultation, documentation, training, and knowledge-sharing.
- Collaborate with the Subsystems Platform Team to create self-service observability and incident management capabilities.
Requirements
- 7+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering with a strong observability and production operations focus.
- Expert hands-on experience with New Relic and strong NRQL proficiency.
- Deep knowledge of structured logging, metrics collection, distributed tracing, dashboards, alerts, SLOs, SLIs, and error budgets.
- Hands-on experience with incident management platforms such as Incident.io, PagerDuty, or OpsGenie, including response workflows, on-call rotations, escalation policies, and post-incident reviews.
- Strong Terraform experience and familiarity with infrastructure-as-code governance patterns.
- Proficiency with PowerShell scripting and strong Azure experience, plus working knowledge of AWS.
- Experience authoring and troubleshooting Azure DevOps pipelines and using Octopus Deploy.
- Comfort working across Windows and Linux server environments and familiarity with Slack for operational workflows.
- Experience with alert noise reduction, observability cost optimization, chaos engineering, game days, failure injection, SQL Server monitoring, FinTech compliance, reliability metrics, Python or Bash, and Jira is desirable.
- Ability to perform hands-on engineering while coaching and mentoring teams in cross-functional Agile/Scrum environments.
Benefits
- Professional development budget
- Flexible in-office collaboration with managers and teams deciding which 10+ days per month to attend
- Team offsites, team bonding activities, and happy hours
- Competitive benefits covering physical and mental healthcare, retirement, family forming, and family support
- Employee giving match
- Mobile phone stipend
- R&R days, generous vacation policy, and parental leave and family planning benefits
- Wellness reimbursement and weekly onsite and virtual programming
- Catered lunches, stocked kitchens, and company events
Categories
DevOpsSite Reliability
About Ripple
Using proven crypto and blockchain technology honed over a decade, Ripple’s enterprise-grade solutions are faster, more transparent, and more cost-effective than traditional financial services. Our customers use these solutions to source crypto, facilitate instant payments, empower their treasury, engage new audiences, lower capital requirements, and drive new revenue. Founded in 2012, Ripple's vision is to enable a world where value moves as seamlessly as information flows today—an Internet of Value. Ripple is the only enterprise blockchain company today with products in commercial use. Ripple’s global payments network includes over 300 customers across 40+ countries and six continents.