3 months ago
Base Salary
$160k - $200k/yr
Responsibilities
- Design and implement New Relic monitoring, alerting, dashboards, and NRQL queries across Azure and AWS.
- Define SLOs, SLIs, and error budgets and coach teams on reliability practices and system health communication.
- Reduce alert noise, improve signal quality, optimize observability costs, and establish logging, metrics, tracing, and dashboard standards.
- Develop and maintain Terraform for monitoring resources and observability infrastructure and enforce IaC governance.
- Author and troubleshoot Azure DevOps pipelines and improve deployment visibility, change tracking, and release hygiene.
- Administer Incident.io, including alert routing, notification workflows, Slack and OpsGenie integrations, and runbooks.
- Establish incident management foundations including postmortems, on-call rotations, escalation policies, severity classifications, and response playbooks.
- Respond to and debrief production incidents, track MTTR and MTTD, and drive continuous improvement.
- Enable engineering teams through workshops, consultation, documentation, training, and knowledge-sharing.
- Collaborate with the Subsystems Platform Team to create self-service observability and incident management capabilities.
Requirements
- 7+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering with a strong observability and production operations focus.
- Expert hands-on experience with New Relic and strong NRQL proficiency.
- Deep knowledge of structured logging, metrics collection, distributed tracing, dashboards, alerts, SLOs, SLIs, and error budgets.
- Hands-on experience with incident management platforms such as Incident.io, PagerDuty, or OpsGenie, including response workflows, on-call rotations, escalation policies, and post-incident reviews.
- Strong Terraform experience and familiarity with infrastructure-as-code governance patterns.
- Proficiency with PowerShell scripting and strong Azure experience, plus working knowledge of AWS.
- Experience authoring and troubleshooting Azure DevOps pipelines and using Octopus Deploy.
- Comfort working across Windows and Linux server environments and familiarity with Slack for operational workflows.
- Experience with alert noise reduction, observability cost optimization, chaos engineering, game days, failure injection, SQL Server monitoring, FinTech compliance, reliability metrics, Python or Bash, and Jira is desirable.
- Ability to perform hands-on engineering while coaching and mentoring teams in cross-functional Agile/Scrum environments.
Benefits
- Professional development budget
- Flexible in-office collaboration with managers and teams deciding which 10+ days per month to attend
- Team offsites, team bonding activities, and happy hours
- Competitive benefits covering physical and mental healthcare, retirement, family forming, and family support
- Employee giving match
- Mobile phone stipend
- R&R days, generous vacation policy, and parental leave and family planning benefits
- Wellness reimbursement and weekly onsite and virtual programming
- Catered lunches, stocked kitchens, and company events
Categories
DevOpsSite Reliability
About Ripple
Ripple builds enterprise blockchain-based payments, treasury, and crypto liquidity products for banks, payment providers, businesses, and governments. It licenses software and provides network services that power cross-border transfers (including On-Demand Liquidity using XRP) and offers a CBDC platform for central banks. Founded in 2012 and headquartered in San Francisco, Ripple serves over 300 customers across more than 40 countries.
