Braze

Senior Site Reliability Engineer

Braze
Apply
4 hours ago
São Paulo, BrazilSenior
H1B Sponsor

Responsibilities

  • Design and operate MongoDB infrastructure to meet enterprise-grade SLAs, including availability, durability, query performance, capacity planning, and sharding.
  • Build proactive MongoDB-specific monitoring and alerting for replication health, oplog lag, lock contention, index hit rates, and related symptoms.
  • Partner with product engineering teams on schema design, index strategies, aggregation pipelines, connection pools, write concerns, and read preferences.
  • Build self-service tooling, automation, backup and restore workflows, point-in-time recovery processes, and operational runbooks.
  • Manage MongoDB cluster provisioning, upgrades, failovers, and decommissions on Kubernetes using the MongoDB Enterprise Kubernetes Operator.
  • Contribute to internal SRE platform tooling in Ruby and/or Go to reduce operational toil.
  • Participate in PagerDuty on-call, lead incident retrospectives, perform root-cause analysis, and implement permanent systemic fixes.

Requirements

  • 5+ years of production experience as a Software Engineer, DevOps Engineer, or Site Reliability Engineer.
  • Hands-on MongoDB expertise including replica sets, sharding, index design, aggregation pipelines, explain plans, and performance tuning under real load.
  • Strong Linux fundamentals and experience operating at the OS level, including disk I/O, memory, networking, and process management.
  • Strong programming skills in one or more of Python, Go, Ruby, or JavaScript, with experience writing automation.
  • Experience with infrastructure-as-code tools such as Terraform or Ansible, or equivalent tools.
  • Experience with Docker and Kubernetes container orchestration.
  • Systems-thinking skills covering interfaces, failure modes, edge cases, and cascading effects across the stack.
  • Experience documenting systems and collaborating asynchronously across global remote teams.
  • Nice-to-have experience running MongoDB at multi-terabyte scale or in a sharded topology.
  • Nice-to-have familiarity with MongoDB Atlas, Ops Manager, or Cloud Manager.
  • Nice-to-have experience with Redis, Kafka, or Postgres and prior database platform or database reliability engineering experience.

Benefits

  • Hybrid working arrangement with a curated in-office employee experience.
  • Competitive compensation that may include equity.
  • Retirement and Employee Stock Purchase Plans.
  • Flexible paid time off.
  • Medical, dental, vision, life, and disability benefit plans.
  • Family services including fertility benefits and equal paid parental leave.
  • Professional development through formal career pathing, learning platforms, and a yearly learning stipend.
  • Opportunities to give back through Volunteer Week and donation matching.
  • Employee Resource Groups and a collaborative, transparent workplace culture.

Categories

DevOpsSite Reliability
Braze

About Braze

1,001-5,000 employees

Braze is the leading customer engagement platform that empowers brands to Be Absolutely Engaging.™ Braze allows any marketer to collect and take action on any amount of data from any source, so they can creatively engage with customers in real time, across channels from one platform. From cross-channel messaging and journey orchestration to Al-powered experimentation and optimization, Braze enables companies to build and maintain absolutely engaging relationships with their customers that foster growth and loyalty. The company has been recognized as a 2024 U.S. News & World Report Best Companies to Work For, 2024 Best Small & Medium Workplaces in Europe by Great Place to Work®, 2024 Fortune Best Workplaces for Women™ by Great Place to Work® and was named a Leader by Gartner® in the 2024 Magic Quadrant™ for Multichannel Marketing Hubs and a Strong Performer in The Forrester Wave™: Email Marketing Service Providers, Q3 2024. Braze is headquartered in New York with 15 offices across AMER, LATAM, EMEA, and APAC. Learn more at braze.com.