Base Salary
$129k - $232k/yr
Responsibilities
- Architect, configure, tune, and operate high-performance NGINX routing, proxying, and Kubernetes ingress controller infrastructure.
- Own and expand automated scaling routines using RED metrics, Horizontal Pod Autoscalers, and customized scaling policies.
- Partner with product engineering teams to translate API feature requirements into highly available, scalable technology stacks.
- Establish SLIs and SLOs, manage error budgets, and conduct systems design, bottleneck profiling, and capacity planning.
- Participate in PagerDuty on-call rotations, update runbooks, and proactively prevent recurring incidents.
- Lead root-cause analyses and blameless retrospectives for availability and performance incidents.
- Build operational tools, scripts, and platform frameworks using Ruby, Go, or equivalent programming languages.
Requirements
- 5+ years of experience as a DevOps or Site Reliability Engineer in a high-scale production environment.
- Deep hands-on experience configuring, troubleshooting, and operating NGINX proxying, routing, and ingress controller layers under heavy traffic.
- In-depth hands-on proficiency with Kubernetes administration, cluster networking, scheduling, orchestration, and container deployment.
- Strong understanding of Linux/Unix internals, including disk I/O, memory allocation, TCP/IP networking, and process management.
- Strong programming or scripting skills; Ruby and Go are preferred, with Python or Java as equivalent alternatives.
- Experience managing infrastructure with Terraform, Ansible, Chef, or similar infrastructure-as-code frameworks.
- Strong understanding of distributed systems design, interfaces, failure modes, edge cases, and cascading effects.
- Excellent documentation practices and comfort collaborating asynchronously across remote-first global engineering teams.
- Preferred familiarity with Redis, Kafka, Postgres, MongoDB, Prometheus, Grafana, Datadog, AWS, GCP, and Azure.
Benefits
- Base pay range of $128,842-$232,200 per year for U.S.-based candidates, with potential equity and other compensation components.
- Retirement and Employee Stock Purchase Plans.
- Flexible paid time off and comprehensive medical, dental, vision, life, and disability benefits.
- Fertility benefits and equal paid parental leave.
- Professional development through formal career pathing, learning platforms, and a yearly learning stipend.
- Hybrid ways of working and a curated in-office employee experience.
- Volunteer Week, donation matching, and Employee Resource Groups.
- Collaborative culture recognized as a Great Place to Work®.
Tech Stack
Categories
About Braze
Braze is the leading customer engagement platform that empowers brands to Be Absolutely Engaging.™ Braze allows any marketer to collect and take action on any amount of data from any source, so they can creatively engage with customers in real time, across channels from one platform. From cross-channel messaging and journey orchestration to Al-powered experimentation and optimization, Braze enables companies to build and maintain absolutely engaging relationships with their customers that foster growth and loyalty. The company has been recognized as a 2024 U.S. News & World Report Best Companies to Work For, 2024 Best Small & Medium Workplaces in Europe by Great Place to Work®, 2024 Fortune Best Workplaces for Women™ by Great Place to Work® and was named a Leader by Gartner® in the 2024 Magic Quadrant™ for Multichannel Marketing Hubs and a Strong Performer in The Forrester Wave™: Email Marketing Service Providers, Q3 2024. Braze is headquartered in New York with 15 offices across AMER, LATAM, EMEA, and APAC. Learn more at braze.com.