Lambda

Senior Site Reliability Engineer - SDN

Lambda
Apply
2 months ago
Bellevue, WA, USA +2 moreSenior
H1B Sponsor

Base Salary

$240k - $312k/yr

Responsibilities

  • Operate and scale Lambda’s multi-tenant cloud networking platform and SDN infrastructure.
  • Operate and improve Kubernetes-based control plane services and dataplane software running on SmartNICs.
  • Develop tooling and automation to reduce operational toil and improve reliability.
  • Collaborate with software, platform, and networking teams on service reliability and deployment workflows.
  • Deploy and maintain network monitoring, observability, and management tools.
  • Improve deployment safety through CI/CD pipelines, GitOps workflows, testing, and progressive rollouts.
  • Drive operational excellence through observability, incident management, capacity planning, postmortems, and on-call participation.

Requirements

  • At least 5 years of experience in Site Reliability Engineering, Production Engineering, or a similar role.
  • Experience operating and supporting large-scale distributed systems in production.
  • Experience with Kubernetes application lifecycle management, upgrades, troubleshooting, and production operations.
  • Experience participating in on-call rotations and incident response.
  • Strong troubleshooting skills across Linux systems, Kubernetes, distributed systems, and networking.
  • Experience with observability platforms, monitoring, alerting, and metrics.
  • Comfort working on the Linux command line and a solid understanding of the Linux networking stack.
  • Experience with multi-datacenter and hybrid cloud environments.
  • Experience automating infrastructure and operational workflows using Python, Ansible, or similar tools.
  • Experience designing and operating CI/CD and GitOps deployment workflows.
  • Preferred experience building and operating SDNs with OpenStack Neutron, OVN, and OVS.
  • Preferred experience with production-scale SDNs in cloud environments, Go and/or Python development, Kubernetes, Helm, Terraform, Ansible, SR-IOV, and DPDK.

Benefits

  • Hybrid work arrangement requiring presence in a San Francisco, San Jose, or Bellevue office four days per week, with Tuesday designated as the work-from-home day.
  • Health, dental, and vision coverage for employees and dependents.
  • Wellness and commuter stipends for select roles.
  • 401(k) plan with a 2% company match for USA employees.
  • Flexible paid time off plan.
  • Cash and equity compensation are offered, with no specific amounts stated.

Tech Stack

Categories

DevOpsSite Reliability
Lambda

About Lambda

501-1,000 employees

The Superintelligence Cloud