Bloomerang

Sr. Software Engineer, Site Reliability

Bloomerang
Apply
2 hours ago
Remote, United StatesSenior

Base Salary

$115k - $150k/yr

Responsibilities

  • Own complex production support escalations and ticket triage through hands-on troubleshooting and resolution.
  • Partner with Software Engineering to identify root causes, reliability risks, defects, and permanent solutions.
  • Lead incident response from triage and mitigation through recovery, root cause analysis, and blameless post-incident reviews.
  • Build observability using metrics, logs, traces, dashboards, actionable alerts, and synthetic monitoring.
  • Define and mature SLIs and SLOs measuring system reliability and customer experience.
  • Reduce operational toil through automation, tooling, process improvements, and permanent fixes.
  • Use AI-assisted tools and source code repositories for triage, troubleshooting, code analysis, automation, and technical investigation.
  • Participate in a rotating on-call schedule, primarily during business hours, with limited after-hours and weekend support.

Requirements

  • Hands-on Site Reliability Engineering experience applying software engineering practices to production reliability and helping establish or mature SRE practices.
  • Strong knowledge of SLIs, SLOs, error budgets, observability, automation, and toil reduction.
  • Experience building monitoring, dashboards, alerts, and telemetry with tools such as Honeycomb, New Relic, Grafana, CloudWatch, or Kibana.
  • Experience with production incident management, root cause analysis, blameless post-incident reviews, and corrective-action follow-through.
  • Strong programming and scripting skills for troubleshooting application code and building automation and operational tooling.
  • Strong SQL and relational database skills for production troubleshooting and safe data correction; PostgreSQL experience is preferred.
  • Strong code literacy and debugging skills, including navigating unfamiliar codebases, understanding application flow, reviewing code and change history, and identifying reliability issues.
  • Experience troubleshooting cloud-hosted applications using source code, logs, APIs, telemetry, event streams, and databases.
  • Comfort navigating application stacks across PHP, .NET, and Node.js; deep expertise in each is not required.
  • Demonstrated experience using AI-assisted tools in day-to-day engineering workflows.
  • Strong communication and collaboration skills across Software Engineering, Product, Support, DevOps, and other technical teams.

Benefits

  • Health, vision, and dental insurance options plus 24/7 access to HealthiestYou healthcare services.
  • 20 PTO days, 3 flex days, 4 optional volunteer days, 12 paid holidays, and paid parental leave.
  • 401k matching.
  • Company-provided equipment shipped to the employee.
  • Permanent, full-time, fully remote work within the U.S. and select Canadian provinces; Indianapolis employees may work from company headquarters.
  • No visa sponsorship or relocation assistance is offered.

Tech Stack

GrafanaKibana.NETNode.jsPHPPostgreSQLSQL

Categories

Site Reliability
Bloomerang

About Bloomerang

201-500 employees

Bloomerang is the Giving Platform built for purpose, trusted by 24,000+ nonprofits to raise more funds, retain more supporters, and create lasting change. By unifying fundraising, CRM, and volunteer management in one easy-to-use platform, Bloomerang gives organizations a complete view of every supporter and the tools to build stronger relationships. Backed by expert support and a team passionate about nonprofit success since 2012, Bloomerang is more than software—it’s a growth partner for missions that matter. Learn more at bloomerang.com.

Contact me