2 months ago
Base Salary
$175k - $300k/yr
Responsibilities
- Design scalable and reliable infrastructure for AI workloads, inference, data pipelines, and agentic workflows.
- Build CI/CD and software development lifecycle tooling to improve developer experience.
- Develop autoscaling based on queue lag, in-flight requests, and latency, including burst capacity and safe drains.
- Evolve Terraform and Helm infrastructure for multi-environment deployments, secrets, policy-as-code, and workload identity.
- Build end-to-end observability across infrastructure, systems, and applications and connect it to the AI SRE agent.
- Improve security and compliance through least privilege, just-in-time access, default-deny egress, auditability, and policy-as-code.
- Operate highly available and resilient cloud and Kubernetes infrastructure, including incident response, chaos testing, and capacity planning.
Requirements
- 7+ years of experience at technically rigorous companies or teams.
- Experience operating cloud- and Kubernetes-native infrastructure and applications at scale with greater than 99.9% availability.
- Hands-on experience with AWS, EKS, Terraform, and Helm.
- Experience designing idempotent systems using patterns such as outbox, deduplication keys, and safe replay.
- Strong debugging skills across infrastructure, compute, network, runtime, storage, and authentication layers.
- Preferred experience with Envoy, Istio, Cilium, eBPF, GPU workload operations, inference servers, token streaming gateways, Python, Rust, TypeScript, data governance, cross-region active-active architectures, GCP, Azure, and Oracle Cloud.
Benefits
- Competitive compensation and startup equity.
- Health insurance and fertility benefits.
- Technology setup stipend and flexible time off.
- In-office snacks, team happy hours and outings, and an annual company offsite.
- Full-time, in-person role based in New York near Madison Square Park, five days per week.