2 months ago
Responsibilities
- Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.
- Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
- Design and improve scalable backend and platform systems, including capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.
- Define and improve dashboards, alerts, logs, traces, SLOs, runbooks, and on-call workflows.
- Build reliable CI/CD, release automation, environment management, and deployment workflows that improve developer productivity and reduce production risk.
- Write maintainable code to automate systems, improve backend services, and create internal tooling.
Requirements
- Experience owning production cloud infrastructure for a high-availability, user-facing platform, including uptime, performance, deployment safety, and cost.
- Deep AWS infrastructure and containerized-systems experience, with Terraform, Kubernetes/EKS, Docker, EC2, CodeBuild, ECR, S3, IAM, load balancers, networking, and secrets management strongly preferred.
- Experience building or operating CI/CD, environment management, release automation, observability, alerting, and incident-response systems.
- Strong backend engineering judgment across service architecture, APIs, databases, asynchronous systems, queues, scaling limits, and production failure modes.
- Ability to write clean, maintainable code and apply software engineering judgment across product architecture, infrastructure, backend systems, and developer workflows.
- Preferred experience with data-heavy, ML/AI, workflow, marketplace, developer-tools, or enterprise platforms.
- Preferred experience designing systems for bursty workloads, long-running jobs, sandboxed execution, distributed workers, or high-concurrency services.
- Preferred experience reducing cloud costs through architecture, autoscaling, workload placement, caching, cleanup systems, or observability.
- Preferred experience building internal platforms or tools that improve developer productivity.
- Technical aptitude, ownership, and learning potential are valued over a specific number of years of experience.
Benefits
- Full-time employment, on-site in the San Francisco Bay Area.
- Relocation and visa support for strong full-time candidates moving to the US.
- 100% covered medical, dental, and vision insurance through Blue Shield of California.
- Lunch and dinner provided when working in the office.
- Company-wide holiday break from Christmas Eve through New Year’s Day, in addition to PTO and paid holidays.
- Equinox membership, 401(k), and commuter benefits.
- Unlimited access to ChatGPT, Claude Code, Cursor, and similar tools.
- Selection process includes two technical interviews and a one-week work trial; applications are reviewed on a rolling basis.
About HUD
The platform to build RL environments. We help improve frontier models across diverse capabilities and help build post-training datasets for doing RFT.
