10 hours ago
Base Salary
$155k - $240k/yr
Responsibilities
- Design scalable, fault-tolerant infrastructure across AWS, GCP, and Azure.
- Operate and improve Kubernetes-based production systems using Helm, Terraform, or related infrastructure tooling.
- Automate operational tasks and infrastructure management with Python or Go.
- Own reliability, performance, and efficiency for core services, including SLOs, error budgets, and on-call operations.
- Lead incident response, post-mortems, root-cause analyses, and reliability improvements.
- Shape long-term platform, observability, cost, and reliability investments.
- Collaborate with product, security, and engineering teams on reliable and scalable system design.
- Use AI-assisted tools to investigate incidents, create infrastructure changes, write runbooks, scaffold tooling, and review pull requests.
Requirements
- 5+ years of experience in infrastructure engineering, DevOps, or a similar role operating large-scale, high-availability production systems.
- Production experience running containerized workloads and real clusters.
- Experience with Helm and Terraform or Pulumi on at least one major cloud provider, preferably AWS.
- Proficiency in Python or Go for automation and tooling.
- Existing daily experience using AI-assisted or agentic tooling, including Claude Code, Droid, Codex, or comparable internal tools.
- Experience with monitoring and logging stacks such as Prometheus, Grafana, and ELK or equivalent.
- Demonstrated ability to identify systemic reliability weaknesses, make tradeoffs, and develop innovative solutions.
- Experience owning at least one infrastructure build end-to-end with measurable outcomes.
- Strong communication, collaboration, problem-solving, autonomy, and end-to-end ownership skills.
- Preferred: software-engineering depth and experience designing and shipping non-trivial production services, libraries, or internal frameworks in Python, Go, or a comparable language.
Benefits
- Hybrid work from WRITER hubs in New York City, San Francisco, Seattle, or London.
- Generous paid time off and company holidays.
- Medical, dental, and vision coverage for employees and families.
- Sixteen weeks of paid parental leave for all parents.
- Fertility and family planning support.
- Early-detection cancer testing through Galleri.
- Flexible spending and dependent care FSA options.
- Health savings account with company contribution for eligible plans.
- Wellness, learning and development, and work-life stipends.
- Company-wide and team off-sites.
- Company stock options and 401(k).
Tech Stack
Categories
DevOpsSite Reliability
About WRITER
WRITER builds an enterprise generative AI platform for creating, deploying, and supervising AI agents grounded in company data and powered by proprietary LLMs. It sells subscriptions and services to large organizations, with customers such as Accenture, Marriott, Uber, Mars, and Vanguard. Founded in 2020 and headquartered in San Francisco, the privately held company raised a Series C in 2024.
