7 months ago
Base Salary
$160k - $350k/yr
Responsibilities
- Own critical production services from design and code review through deployment, operation, and incident response.
- Profile, benchmark, and rewrite performance-critical paths to eliminate bottlenecks.
- Lead incident response and convert post-mortem findings into code and architectural improvements.
- Build observability frameworks, instrumentation, alerting logic, and debugging tools.
- Define and enforce SLOs across platform services and create accountability feedback loops.
- Own capacity planning and cost efficiency through growth modeling, infrastructure right-sizing, and automation.
- Build and improve well-tested internal platforms and deployment tooling.
- Own CI/CD systems and enable safe, efficient software delivery.
- Embed with product engineering teams, contribute to production codebases, and co-design reliable systems.
- Partner on infrastructure security through threat modeling, hardening, and automated compliance tooling.
Requirements
- At least five years of software development experience shipping and maintaining production services.
- Production-grade proficiency in at least one of Go, Python, C++, or Rust.
- Experience as a Production Engineer, SRE, or software engineer with a deep infrastructure focus and end-to-end service ownership.
- Deep understanding of distributed systems.
- Container orchestration expertise and experience debugging complex distributed failures in production.
- Working knowledge of operating-system-level concepts.
- Fluency with cloud platforms, preferably AWS.
- Experience building and maintaining observability stacks.
- Strong CI/CD pipeline expertise and experience improving developer velocity safely.
- Experience at a company with a Production Engineering or software-focused SRE culture is a strong plus.
- Experience building platforms for AI/ML workloads or high-throughput document processing pipelines is a plus.
Benefits
- Unlimited PTO
- Medical, dental, and vision insurance
- 401(k)
- Catered lunch daily
- DoorDash dinner credit when staying late
- Three months of parental leave for non-birthing parents and four months for birthing parents
- $15,000 lifetime fertility benefit
- Competitive new-hire equity grant
- On-site role
About Hebbia
Hebbia builds a generative-AI research and document analysis platform for finance and professional services, used to surface insights across filings, research, and internal data and to automate analyst workflows. Its flagship product, Matrix, is sold as enterprise software to investment banks, asset managers, private equity firms, and corporate finance teams. Founded in 2020 and headquartered in New York, the company is privately held and counts major Wall Street firms among its customers.
