Axon

Senior Site Reliability Engineer I

Axon
Apply
2 days ago

Base Salary

$134k - $215k/yr

Responsibilities

  • Own and evolve distributed tracing infrastructure using OpenTelemetry and Jaeger and drive instrumentation adoption across Axon’s service-oriented architecture.
  • Build and operate the Grafana Loki and Alloy log aggregation platform, including expanding use cases beyond Kubernetes event logs.
  • Maintain Cortex, Prometheus, and Grafana metrics infrastructure supporting alerting, dashboards, and SLO tracking.
  • Create internal tooling and automation for self-service observability, including toolkit commands, on-call helpers, runbook generation, and dashboard scaffolding.
  • Manage observability infrastructure as code with Terraform, CDK, ArgoCD, and Helm while addressing capacity, cybersecurity, compliance, and on-call needs.
  • Partner with engineering teams to establish instrumentation standards and meaningful service-level objectives.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or an equivalent highly technical field.
  • At least 7 years of experience in SRE, platform engineering, or infrastructure engineering.
  • Experience with agentic AI tooling or building LLM-powered developer tools.
  • Strong Linux systems fundamentals and experience working in Kubernetes-based environments.
  • Hands-on experience with one or more LGTM components, including Loki, Grafana, Tempo, Jaeger, Mimir, or Cortex.
  • Experience with infrastructure as code; Terraform is strongly preferred and CDK experience is a plus.
  • Experience with Golang, Python, or Java.
  • United States citizenship and ability to obtain CJIS clearance for full U.S. production access.
  • Preferred qualifications include distributed tracing operations, OpenTelemetry instrumentation and pipelines, GitOps workflows, high-volume systems with formal SLA requirements, and debugging complex multi-service distributed systems.

Benefits

  • Competitive salary and 401(k) with employer match.
  • Discretionary time off and paid parental leave for all employees.
  • Medical, dental, and vision plans.
  • Fitness programs and emotional and development programs.
  • Seattle-based hybrid schedule with onsite work Tuesday through Friday and remote flexibility on Mondays.
  • Participation in an on-call rotation.

Tech Stack

GoGrafanaHelmJavaKubernetesLinuxPrometheusPythonSplunkTerraform

Categories

DevOpsSite Reliability
Contact me