3 days ago
Base Salary
$147k - $224k/yr
Responsibilities
- Manage Kubernetes clusters across multiple environments and regions.
- Own infrastructure as code for all resources.
- Maintain and improve CI/CD pipelines and GitOps-based deployments.
- Maintain and optimize real-time data pipelines processing billions of events per day across distributed queues and stream processors.
- Build monitoring, alerting, and observability systems.
- Debug production issues across services.
- Manage cloud costs and capacity planning.
- Own the company’s infrastructure while working closely with a small engineering team.
Requirements
- 5–8 years of experience in a production DevOps or SRE role.
- Proven experience designing and operating large-scale distributed systems, including API design, reliability, and performance at scale.
- Strong Kubernetes experience in a managed cloud environment.
- Proficiency with Terraform or a similar infrastructure-as-code tool.
- Experience with GitOps-based deployment workflows.
- Experience building or maintaining observability stacks covering logging, metrics, and alerting.
- Experience handling production incidents calmly and methodically.
- Nice-to-have experience with multi-region deployments, search infrastructure, streaming or warehousing data pipelines, and large-scale proxy or networking infrastructure.
- Applicants must be authorized to work in the country where they apply and provide proof of employment eligibility.
Benefits
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to a 4% company match and immediate vesting.
- 20 weeks of paid parental leave for primary caregivers and 12 weeks for secondary caregivers.
- Up to $85 per month in mobile and internet reimbursement.
- Company-paid short-term disability, long-term disability, and life insurance.
- Career growth and learning opportunities, flexibility and ownership, and a collaborative international work environment.
- In-office work is required.
Tech Stack
Categories
DevOpsSite Reliability
About Nebius
The Nebius AI Cloud brings powerful full-stack infrastructure for AI developers and practitioners across startups, enterprises and science institutes to build and deploy generative AI applications and rapidly deliver scientific breakthroughs by training and running ML models within a secure, high-performance, and cost-optimized cloud environment.
