3 months ago
Base Salary
$160k - $200k/yr
Responsibilities
- Design and implement New Relic monitoring, alerting, dashboards, and NRQL queries across Azure and AWS
- Define SLOs, SLIs, and error budgets and coach teams on reliability management
- Reduce alert noise, improve signal quality, and optimize observability costs through log-ingestion and configuration governance
- Develop and maintain Terraform for monitoring resources, alert configurations, and observability infrastructure
- Establish auditable infrastructure-as-code governance standards across teams
- Author and troubleshoot Azure DevOps pipelines and improve deployment visibility and release hygiene
- Administer Incident.IO, including alert routing, notification workflows, Slack and OpsGenie integrations, and runbooks
- Establish postmortem processes, on-call rotations, escalation policies, incident severity classifications, and response playbooks
- Respond to and debrief production incidents while tracking MTTR, MTTD, incident frequency, availability, and error budgets
- Enable engineering teams through workshops, consultation, documentation, training, and knowledge-sharing
- Partner with the Subsystems Platform Team to create self-service observability and incident-management capabilities
Requirements
- 7+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering focused on observability and production operations
- Hands-on ability to deliver engineering work while coaching and mentoring teams
- Expertise with New Relic and NRQL, including APM, Infrastructure, Logs, Synthetics, and Alerts
- Deep understanding of structured logging, metrics collection using RED/USE methods, distributed tracing, dashboards, alerts, SLOs, SLIs, and error budgets
- Hands-on experience with incident-management platforms such as Incident.IO, PagerDuty, or OpsGenie
- Experience designing incident workflows, on-call rotations, escalation policies, post-incident reviews, and production troubleshooting
- Strong Terraform experience and familiarity with infrastructure-as-code governance
- Proficiency with PowerShell scripting
- Strong Azure experience, including App Services, Virtual Machines, Azure SQL, networking, and monitoring, plus working knowledge of AWS
- Experience authoring and troubleshooting Azure DevOps pipelines and using Octopus Deploy
- Comfort working across Windows and Linux server environments
- Familiarity with Slack for operational workflows and incident communication
- Experience with alert-noise reduction, observability cost optimization, chaos engineering, game days, failure injection, SQL Server monitoring, FinTech compliance, reliability metrics, Python or Bash, and Jira is desirable
Benefits
- Professional development budget and learning opportunities
- Flexible in-office collaboration with teams deciding which 10+ days per month to work in the office
- Team offsites, team bonding activities, happy hours, and bi-weekly all-company meetings
- Healthcare, retirement, family-forming, and family-support benefits
- Employee giving match and mobile phone stipend
- R&R days, generous vacation policy, wellness reimbursement, and onsite and virtual wellness programming
- Parental leave and family-planning benefits
- Catered lunches, stocked kitchens, snacks, beverages, and company events
- Full-time employee benefits apply to this position
Categories
DevOpsSite Reliability
About Ripple
Ripple builds enterprise blockchain-based payments, treasury, and crypto liquidity products for banks, payment providers, businesses, and governments. It licenses software and provides network services that power cross-border transfers (including On-Demand Liquidity using XRP) and offers a CBDC platform for central banks. Founded in 2012 and headquartered in San Francisco, Ripple serves over 300 customers across more than 40 countries.
