3 days ago
Remote, United StatesMid Level
Base Salary
$130k - $180k/yr
Responsibilities
- Ensure fault tolerance, scalability, and uninterrupted operations for infrastructure services.
- Use technology to solve infrastructure problems and optimize system performance.
- Implement and improve CI/CD processes.
- Troubleshoot complex hardware, software, and networking issues.
- Support monitoring, testing, asset tracking, repair tracking, and server production across data-center infrastructure.
- Collaborate with globally distributed engineering and operations teams.
Requirements
- Proficiency in Linux systems.
- Expertise in Python and Bash scripting for automation.
- Demonstrated ability to troubleshoot complex hardware, software, and networking problems.
- Strong analytical and problem-solving skills focused on system performance optimization.
- Working proficiency in English.
- Experience with backend development is a preferred qualification.
- Experience designing, developing, and running high-load distributed systems is a preferred qualification.
Benefits
- Primarily remote work with occasional required travel to data centers.
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to a 4% company match and immediate vesting.
- Paid parental leave of 20 weeks for primary caregivers and 12 weeks for secondary caregivers.
- Remote work reimbursement of up to $85 per month for mobile and internet.
- Company-paid short-term disability, long-term disability, and life insurance.
- Career growth and learning opportunities in an international environment.
About Nebius
The Nebius AI Cloud brings powerful full-stack infrastructure for AI developers and practitioners across startups, enterprises and science institutes to build and deploy generative AI applications and rapidly deliver scientific breakthroughs by training and running ML models within a secure, high-performance, and cost-optimized cloud environment.
