2 months ago
Base Salary
$192k - $240k/yr
Responsibilities
- Design, build, and maintain release infrastructure powering deployment pipelines and incident workflows.
- Drive technical strategy and architecture for scalable, reliable, and secure release and observability systems.
- Improve the end-to-end release process from code merge through production to reduce risk and cycle time.
- Build and evolve observability and incident response tooling for rapid detection, triage, and resolution.
- Identify and mitigate performance, reliability, and security risks in the release and infrastructure stack.
- Define, instrument, and monitor release engineering metrics such as deployment frequency, change failure rate, and MTTR.
- Partner with infrastructure and product teams to debug production issues and deliver long-term fixes.
- Champion release engineering, reliability, and operational excellence practices across the organization.
- Mentor engineers through technical guidance and code reviews.
- Evaluate emerging release engineering, observability, and SRE tools and practices for adoption.
Requirements
- At least 7 years of professional experience designing, building, and operating backend or infrastructure systems in production.
- Strong proficiency in backend programming languages such as Go, Java, Kotlin, or Python, with a focus on reliability and performance.
- Hands-on experience with CI/CD and release pipelines, including build, test, and deployment automation.
- Experience architecting and operating scalable, highly available distributed systems on AWS, GCP, or Azure.
- Deep familiarity with Docker, Kubernetes, Terraform, and CloudFormation.
- Experience designing and maintaining observability tooling for metrics, logs, and tracing and integrating it with incident response workflows.
- Strong understanding of reliability and SRE practices, including SLIs, SLOs, error budgets, and incident management.
- Experience designing and optimizing SQL and/or NoSQL data storage systems for operational and observability use cases.
- A proven record of improving release processes through risk reduction, increased deployment frequency, or automated rollbacks.
- Ability to work cross-functionally to debug complex production issues and ship changes safely.
- Strong communication and collaboration skills, including writing clear design documents and driving technical decisions across teams.
Benefits
- Based in the San Francisco office in a hybrid arrangement requiring at least three coordinated in-office days per week on Monday, Wednesday, and Thursday.
- Up to four weeks per year of fully remote work.
Tech Stack
Categories
DevOpsSite Reliability
About Brex
Brex is the intelligent finance platform built for speed and control, empowering founders and finance teams to spend smarter and move faster globally with intuitive corporate cards, banking, expenses, and travel. Over 35,000 companies run on Brex, including DoorDash, ServiceTitan, Wiz, and Five Guys. Brex LLC is a wholly owned subsidiary of Capital One, N.A.