5 days ago
Remote, United StatesSenior
Base Salary
$168k - $185k/yr
Responsibilities
- Design, build, and operate shared cloud infrastructure using AWS, Kubernetes, Terraform, Databricks, Cloudflare, and related cloud-native technologies.
- Deliver SRE and DevOps initiatives that improve reliability, scalability, observability, deployment safety, and operational readiness.
- Build reusable infrastructure modules, automation, and self-service workflows that improve developer experience and reduce manual work.
- Define and implement service-level indicators, service-level objectives, monitoring, alerting, and error-budget practices.
- Participate in incident response, post-incident reviews, runbook development, and corrective actions.
- Improve disaster-recovery readiness through recovery planning, automation, testing, and remediation.
- Improve CI/CD workflows and infrastructure delivery processes.
- Partner with application, Data Platform, and Data Engineering teams to support workload operations and integrations.
- Improve cloud efficiency through architecture, capacity planning, Kubernetes resource optimization, cost visibility, and automation.
- Contribute to technical standards, architecture decisions, documentation, and a sustainable 24/7 operating model.
Requirements
- 4+ years of relevant software, infrastructure, SRE, platform engineering, or DevOps experience.
- Hands-on production experience with AWS, Kubernetes, Datadog, Databricks, and Terraform.
- Ability to write reliable, maintainable software and automation across infrastructure, systems, and application boundaries.
- Experience operating business-critical production systems, including observability, incident response, operational readiness, and root-cause analysis.
- Experience improving CI/CD systems, infrastructure as code, deployment workflows, or internal developer platforms.
- Understanding of SLIs, SLOs, error budgets, capacity planning, and disaster-recovery objectives.
- Interest in AI-native and data-intensive applications and experience or willingness to work with Databricks, OpenAI, Anthropic, and open-source AI tools.
- Ability to independently own projects, make technical tradeoffs, communicate clearly, and collaborate across teams.
- Comfort participating in an on-call or incident-response rotation.
- Experience with data-intensive SaaS, analytics, big data analytics, or rapidly scaling technology environments is preferred.
- Bachelor’s degree or equivalent practical experience.
Benefits
- Fully remote work within the United States.
- East Coast working hours are expected.
- Flexible work hours and flexible vacation.
- Generous 401(k) match, parental leave, team events, wellness budget, learning reimbursement, and other benefits.
- Annual base salary range of $168,000–$185,000 plus equity.
- Remote work outside company offices may be subject to New York State tax withholding.
Tech Stack
Categories
DevOpsSite Reliability
