Braze

Senior Site Reliability Engineer I

Braze
Apply
3 hours ago
Vancouver, CanadaSenior
H1B Sponsor

Responsibilities

  • Configure, tune, and operate high-performance NGINX routing, proxying, and Kubernetes ingress controller layers.
  • Own and expand automated scaling routines using RED metrics, Horizontal Pod Autoscalers, and customized scaling policies.
  • Partner with product engineering teams to translate API feature requirements into resilient, highly available, and scalable technology stacks.
  • Establish SLIs and SLOs, manage error budgets, and support enterprise-grade service levels through systems design and capacity planning.
  • Participate in PagerDuty on-call rotations, improve runbooks, and prevent recurring incidents.
  • Lead root-cause analyses and blameless retrospectives and turn operational learnings into permanent improvements.

Requirements

  • 5+ years of experience as a DevOps or Site Reliability Engineer in a high-scale production environment.
  • Deep hands-on experience configuring, troubleshooting, and operating NGINX proxying, routing, and ingress controller layers under heavy traffic.
  • In-depth hands-on Kubernetes administration experience, including cluster networking, orchestration, scheduling, and container deployment.
  • Strong Linux/Unix knowledge covering disk I/O, memory allocation, TCP/IP networking, and process management.
  • Strong programming or scripting skills; Ruby and/or Go are preferred, with Python or Java accepted as equivalent languages.
  • Experience managing infrastructure with Terraform, Ansible, Chef, or similar infrastructure-as-code frameworks.
  • Strong understanding of distributed systems design, interfaces, failure modes, edge cases, and cascading effects.
  • Strong documentation and collaboration skills for asynchronous work across remote-first global engineering teams.
  • Familiarity with Redis, Kafka, Postgres, or MongoDB is preferred.
  • Experience with Prometheus, Grafana, Datadog, or similar monitoring and observability tools is preferred.
  • Practical experience with AWS, GCP, or Azure is preferred.

Benefits

  • Location-dependent comprehensive benefits, including medical, dental, vision, life, and disability coverage.
  • Retirement and Employee Stock Purchase Plans.
  • Flexible paid time off and equal paid parental leave with family services including fertility benefits.
  • Professional development through formal career pathing, learning platforms, and a yearly learning stipend.
  • Hybrid ways of working and a curated in-office employee experience.
  • Volunteer Week, donation matching, and Employee Resource Groups.
  • Equity eligibility through restricted stock units (RSUs).

Tech Stack

AnsibleApache KafkaAWSAzureChefDatadogGoGoogle Cloud PlatformGrafanaJavaKubernetesLinuxMongoDBPostgreSQLPrometheusPythonRedisRubyRuby on RailsTerraform

Categories

Site Reliability
Braze

About Braze

1,001-5,000 employees

Braze is the leading customer engagement platform that empowers brands to Be Absolutely Engaging.™ Braze allows any marketer to collect and take action on any amount of data from any source, so they can creatively engage with customers in real time, across channels from one platform. From cross-channel messaging and journey orchestration to Al-powered experimentation and optimization, Braze enables companies to build and maintain absolutely engaging relationships with their customers that foster growth and loyalty. The company has been recognized as a 2024 U.S. News & World Report Best Companies to Work For, 2024 Best Small & Medium Workplaces in Europe by Great Place to Work®, 2024 Fortune Best Workplaces for Women™ by Great Place to Work® and was named a Leader by Gartner® in the 2024 Magic Quadrant™ for Multichannel Marketing Hubs and a Strong Performer in The Forrester Wave™: Email Marketing Service Providers, Q3 2024. Braze is headquartered in New York with 15 offices across AMER, LATAM, EMEA, and APAC. Learn more at braze.com.

Contact me