Forward Networks

Site Reliability Engineer

Forward Networks
Apply
2 months ago
Santa Clara, CA, USASenior
H1B Sponsor

Base Salary

$230k - $250k/yr

Responsibilities

  • Define and drive SRE practices, including SLOs, SLIs, error budgets, and engineering frameworks.
  • Drive reliability and operational excellence for Forward’s SaaS platform.
  • Build and maintain observability infrastructure for logging, metrics, tracing, and alerting.
  • Lead incident response through on-call rotations, runbooks, post-mortems, and corrective follow-through.
  • Partner with engineering teams on capacity planning, load testing, chaos engineering, and production readiness reviews.
  • Help define and build the SRE team as the company scales.

Requirements

  • 6+ years of experience in site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environment.
  • Proven experience building or significantly maturing an SRE function.
  • Strong networking fundamentals, including TCP/IP, DNS, routing, switching, firewalls, and load balancing.
  • Hands-on production experience with Kubernetes and container orchestration.
  • Deep proficiency with observability tools such as Prometheus, Grafana, Datadog, or Splunk.
  • Strong scripting and automation skills in Python, Bash, or similar technologies.
  • Experience with AWS, GCP, or Azure and infrastructure as code using Terraform, Ansible, or equivalent.
  • Experience owning and improving incident-response processes, including blameless post-mortems and SLO-driven reliability improvements.
  • Clear communication skills with engineering teams, non-technical stakeholders, customer-facing teams, and executives.
  • Experience supporting enterprise or federal government customers with high-availability requirements is preferred.
  • Experience as a foundational or early SRE hire at a growth-stage company is preferred.

Benefits

  • Base pay range of $230,000–$250,000 per year.

Tech Stack

AnsibleAWSAzureBashDatadogGoogle Cloud PlatformGrafanaKubernetesPrometheusPythonSplunkTerraform

Categories

DevOpsSite Reliability
Forward Networks

About Forward Networks

201-500 employees

The future of network operations is network modeling. Forward's flagship platform, Forward Enterprise, gives users a mathematically accurate network digital twin. Forward enables perfect network visibility, full path analysis, and security policy verification, freeing up time and saving you money. Get a demo today: www.forwardnetworks.com

Contact me