Base Salary
$129k - $232k/yr
Responsibilities
- Configure, tune, and operate high-performance NGINX routing, proxying, and ingress controller layers.
- Own and expand automated scaling routines for high-throughput API services using RED metrics, Horizontal Pod Autoscalers, and customized scaling policies.
- Partner with product engineering teams to translate API feature requirements into resilient, highly available, and scalable technology stacks.
- Establish SLIs and SLOs and help teams manage error budgets.
- Conduct systems design, bottleneck profiling, and capacity planning for enterprise-grade service levels.
- Participate in PagerDuty on-call rotations, improve runbooks, and prevent recurring incidents.
- Lead root-cause analyses and blameless retrospectives and implement permanent system improvements.
Requirements
- 5+ years of experience as a DevOps or Site Reliability Engineer in a high-scale production environment.
- Deep hands-on experience configuring, troubleshooting, and operating NGINX proxying, routing, and ingress controller layers under heavy traffic.
- In-depth hands-on proficiency with Kubernetes administration, cluster networking, scheduling, orchestration, and container deployment.
- Excellent understanding of Linux/Unix internals, including disk I/O, memory allocation, TCP/IP networking, and process management.
- Strong programming or scripting skills; Ruby and Go are preferred, with Python or Java as equivalent alternatives.
- Experience managing infrastructure with Terraform, Ansible, Chef, or similar infrastructure-as-code frameworks.
- Strong understanding of distributed systems design, interfaces, failure modes, edge cases, and cascading effects.
- Excellent documentation practices and comfort collaborating asynchronously across remote-first global engineering teams.
- Familiarity with Redis, Kafka, Postgres, MongoDB, Prometheus, Grafana, Datadog, AWS, GCP, or Azure is preferred.
Benefits
- Base pay for U.S. candidates is listed at $128,842–$232,200 per year, with additional OTE and potential equity; compensation figures are excluded from benefit classification.
- Retirement and Employee Stock Purchase Plans.
- Flexible paid time off.
- Medical, dental, vision, life, and disability benefits.
- Fertility benefits and equal paid parental leave.
- Formal career pathing, learning platforms, professional development, and a yearly learning stipend.
- Curated in-office employee experience and hybrid ways of working.
- Volunteer Week, donation matching, and Employee Resource Groups.
- Collaborative culture recognized as a Great Place to Work®.
Tech Stack
Categories
About Braze
Braze is the leading customer engagement platform that empowers brands to Be Absolutely Engaging.™ Braze allows any marketer to collect and take action on any amount of data from any source, so they can creatively engage with customers in real time, across channels from one platform. From cross-channel messaging and journey orchestration to Al-powered experimentation and optimization, Braze enables companies to build and maintain absolutely engaging relationships with their customers that foster growth and loyalty. The company has been recognized as a 2024 U.S. News & World Report Best Companies to Work For, 2024 Best Small & Medium Workplaces in Europe by Great Place to Work®, 2024 Fortune Best Workplaces for Women™ by Great Place to Work® and was named a Leader by Gartner® in the 2024 Magic Quadrant™ for Multichannel Marketing Hubs and a Strong Performer in The Forrester Wave™: Email Marketing Service Providers, Q3 2024. Braze is headquartered in New York with 15 offices across AMER, LATAM, EMEA, and APAC. Learn more at braze.com.