17 days ago
Base Salary
$307k - $352k/yr
Responsibilities
- Design and implement reusable, self-service infrastructure on AWS using infrastructure-as-code patterns.
- Own production systems through reliability, security, capacity, cost management, upgrades, incidents, and recovery.
- Provide critical services with SLOs, actionable alerts, dashboards, runbooks, and tested recovery paths.
- Build paved paths, self-service tools, and automated workflows that improve developer delivery speed and reduce toil.
- Set technical direction and carry ambiguous infrastructure work from problem framing through implementation, rollout, and production ownership.
- Partner with Infrastructure, Security, Data, and Product Engineering to build security and compliance controls into engineering workflows.
- Review code and designs, establish reusable standards and documentation, and lead through incidents and unfamiliar failure modes.
Requirements
- 7+ years of infrastructure, platform, or software engineering experience, or an equivalent record of Staff-level technical impact.
- Sustained ownership and operation of a consequential production system, including incident response and post-incident improvement.
- Strong coding skills in TypeScript, Python, Go, Ruby, or a similar language, with the ability to build production software.
- Production experience with AWS, infrastructure as code, containers, CI/CD, observability, and related security boundaries.
- Depth in at least one infrastructure domain with the judgment to work across adjacent infrastructure areas.
- Experience building platform capabilities that engineers adopt because they solve real problems and are easier to use than one-off alternatives.
- Clear, direct communication and the ability to explain technical decisions, evidence, and tradeoffs.
- Preferred experience includes observability, developer productivity, cloud compute, networking, storage, data infrastructure, or security.
- Preferred experience includes Pulumi and TypeScript on AWS, including ECS or Fargate, Kubernetes, RDS, S3, Lambda, VPC, and IAM.
- Experience operating an early-stage platform through rapid growth in fintech or another regulated environment is preferred.
