20 days ago
Remote, SpainSenior
Responsibilities
- Develop, implement, maintain, and troubleshoot cloud and AI infrastructure based on open source software.
- Deploy AI infrastructure on Kubernetes and NVIDIA-certified hardware according to engineering designs.
- Optimize infrastructure performance, reliability, scalability, and security across networking, storage, Linux, and Kubernetes environments.
- Troubleshoot and resolve complex technical issues involving hardware and software infrastructure.
- Design and implement AI-driven automation across the DevOps lifecycle, including code development and maintenance.
- Collaborate with stakeholders and distributed international teams to gather requirements, define technical strategies, and improve processes.
- Lead technical tasks, participate in code reviews, and maintain high-quality software and services.
- Mentor team members and Mirantis customers and facilitate customer knowledge transfer during delivery.
- Stay current with cloud operations and development trends and best practices.
- Travel internationally up to 25% when needed.
Requirements
- At least 5 years of professional experience in DevOps, software development, or a similar role.
- Bachelor’s degree in Computer Science or a related field, or equivalent experience.
- Strong experience with cloud and infrastructure technologies, including Kubernetes and/or OpenStack.
- Experience with high-performance data center processing, networking, and storage.
- Exposure to Golang and working knowledge of Python, JavaScript, or other programming languages.
- Strong knowledge of distributed systems, microservices architecture, and CI/CD pipelines.
- Strong debugging and performance-optimization skills across networking, storage, Linux, and Kubernetes, with attention to security.
- Ability to lead technical tasks, work independently, and collaborate with geographically distributed teams and customers.
- Excellent written and spoken English and strong customer-facing communication skills.
- Preferred qualifications include network or storage architecture experience; high-performance computing or GPU infrastructure experience; GPU scheduling, MIG/vGPU, RDMA/RoCE, InfiniBand, NVLink, DCGM, GPU driver or firmware lifecycle, or NVIDIA AI Enterprise experience; open source contributions or conference presentations; and experience with Rancher, OpenShift, or VMware.
Benefits
- Professional development and training.
- Opportunities to attend conferences and working groups.
- Company outings, happy hours, hackathons, and tech talks.
- Work with passionate colleagues and Fortune 500 and Global 2000 customers on open-source cloud infrastructure innovation.
- Competitive compensation package with a strong benefits plan.
- International travel may be required up to 25%.
Tech Stack
Categories
DevOpsSite Reliability
