20 days ago
Remote, Czechia or Prague, CzechiaSenior
Responsibilities
- Develop, implement, maintain, and troubleshoot cloud and AI infrastructure solutions based on open source software.
- Deploy AI infrastructure on NVIDIA-certified hardware and implement architecture and designs produced by the engineering team.
- Optimize system performance, reliability, scalability, and security across container infrastructure.
- Troubleshoot, debug, and resolve complex issues involving networking, storage, Linux, Kubernetes, and related hardware and software.
- Design and implement AI-driven automation across the DevOps lifecycle, including code development and maintenance.
- Collaborate with stakeholders to gather and refine technical requirements and define technical strategies.
- Participate in code reviews and maintain high-quality software and services.
- Work with geographically distributed international teams on technical challenges and process improvements.
- Mentor team members and Mirantis customers and facilitate knowledge transfer during delivery phases.
- Lead technical tasks, make independent technical judgments, and communicate directly with customers.
- Stay current with cloud operations and development trends and best practices.
Requirements
- At least 5 years of professional experience in DevOps, software development, or a similar role.
- Bachelor's degree in Computer Science or a related field, or equivalent experience.
- Strong experience with cloud and infrastructure technologies, including Kubernetes and/or OpenStack.
- Experience with high-performance data center processing, networking, and storage.
- Exposure to Golang and working knowledge of Python and JavaScript.
- Strong knowledge of distributed systems, microservices architecture, and CI/CD pipelines.
- Strong troubleshooting and debugging skills across networking, storage, Linux, and Kubernetes, with attention to performance optimization and security.
- Demonstrated ability to lead technical tasks and collaborate effectively with diverse teams.
- Excellent written and spoken English and strong customer-facing communication skills.
- Ability to work independently with limited day-to-day oversight and make judgment calls directly with customers.
- Ability to travel up to 25%, including internationally if needed.
- Preferred experience with network or storage architecture, high-performance computing or GPU infrastructure, GPU scheduling, MIG/vGPU, RDMA/RoCE, InfiniBand, NVLink, DCGM health-checking, GPU driver or firmware lifecycle, or NVIDIA AI Enterprise.
- Preferred presence in the open source community through upstream contributions or conference presentations.
- Preferred experience with Rancher, OpenShift, or VMware.
Benefits
- Professional development and training.
- Opportunities to attend conferences and working groups.
- Company outings, happy hours, hackathons, and tech talks.
- Competitive compensation package with a strong benefits plan.
- Opportunity to work with passionate colleagues and Fortune 500 and Global 2000 customers on open source cloud technologies.
Tech Stack
Categories
Site Reliability
