Responsibilities
- Architect, configure, tune, and operate high-performance NGINX routing, proxying, and Kubernetes ingress infrastructure.
- Own and expand automated scaling routines using RED metrics, Horizontal Pod Autoscalers, and customized scaling policies.
- Partner with product engineering teams to translate API requirements into resilient, highly available, and scalable technology stacks.
- Establish SLIs and SLOs, manage error budgets, and conduct systems design, bottleneck profiling, and capacity planning.
- Participate in PagerDuty on-call rotations, update runbooks, lead root-cause analysis, and conduct blameless retrospectives.
- Translate operational learnings into permanent improvements to availability, performance, and reliability.
Requirements
- 5+ years of experience as a DevOps or Site Reliability Engineer in a high-scale production environment.
- Deep hands-on experience configuring, troubleshooting, and operating NGINX proxying, routing, and ingress controller layers under heavy traffic.
- In-depth hands-on Kubernetes administration experience, including cluster networking, orchestration, scheduling, and container deployment.
- Strong OS-level understanding of Linux or Unix internals, including disk I/O, memory allocation, TCP/IP networking, and process management.
- Strong programming or scripting skills; Ruby and/or Go are preferred, with Python or Java as equivalent examples.
- Experience managing infrastructure with Terraform, Ansible, Chef, or similar infrastructure-as-code frameworks.
- Strong systems-design understanding across interfaces, boundaries, failure modes, edge cases, and cascading effects in distributed architectures.
- Excellent documentation practices and comfort collaborating asynchronously across remote-first global engineering teams.
- Familiarity with Redis, Kafka, Postgres, MongoDB, Prometheus, Grafana, Datadog, AWS, GCP, or Azure is preferred.
Benefits
- Compensation may include equity, along with retirement and employee stock purchase plans.
- Flexible paid time off and comprehensive medical, dental, vision, life, and disability benefits are offered.
- Family services include fertility benefits and equal paid parental leave.
- Professional development includes formal career pathing, learning platforms, and a yearly learning stipend.
- The company offers a curated in-office employee experience and hybrid ways of working.
- Employees can participate in Volunteer Week, donation matching, and Employee Resource Groups.
- Benefits vary by location, and the role includes a PagerDuty on-call rotation.
Tech Stack
Categories
About Braze
Braze is the leading customer engagement platform that empowers brands to Be Absolutely Engaging.™ Braze allows any marketer to collect and take action on any amount of data from any source, so they can creatively engage with customers in real time, across channels from one platform. From cross-channel messaging and journey orchestration to Al-powered experimentation and optimization, Braze enables companies to build and maintain absolutely engaging relationships with their customers that foster growth and loyalty. The company has been recognized as a 2024 U.S. News & World Report Best Companies to Work For, 2024 Best Small & Medium Workplaces in Europe by Great Place to Work®, 2024 Fortune Best Workplaces for Women™ by Great Place to Work® and was named a Leader by Gartner® in the 2024 Magic Quadrant™ for Multichannel Marketing Hubs and a Strong Performer in The Forrester Wave™: Email Marketing Service Providers, Q3 2024. Braze is headquartered in New York with 15 offices across AMER, LATAM, EMEA, and APAC. Learn more at braze.com.