Mirantis

Senior Site Reliability Engineer (SRE)

Mirantis
Apply
20 days ago
Remote, BulgariaSenior

Responsibilities

  • Develop, implement, maintain, deploy, and troubleshoot cloud and AI infrastructure based on open source software.
  • Optimize the performance, reliability, scalability, and security of container infrastructure.
  • Troubleshoot and resolve complex issues across networking, storage, Linux, Kubernetes, and hardware and software systems.
  • Design and implement AI-driven automation across the DevOps lifecycle.
  • Collaborate with stakeholders to gather technical requirements and define technical strategies.
  • Participate in code reviews and maintain high-quality software and services.
  • Lead technical tasks, mentor team members, and facilitate knowledge transfer to customers.
  • Work with geographically distributed international teams and support customer delivery phases.
  • Travel internationally up to 25% if needed.

Requirements

  • At least 5 years of professional experience in DevOps, software development, or a similar role.
  • Bachelor's degree in Computer Science or a related field, or equivalent experience.
  • Strong experience with cloud and infrastructure technologies, including Kubernetes and/or OpenStack.
  • Experience with high-performance data center processing, networking, and storage.
  • Exposure to Golang and working knowledge of Python and JavaScript.
  • Strong knowledge of distributed systems, microservices architecture, and CI/CD pipelines.
  • Strong debugging and problem-solving skills across networking, storage, Linux, and Kubernetes, with attention to performance optimization and security.
  • Ability to lead technical tasks, collaborate with diverse teams, make independent decisions, and work directly with customers.
  • Excellent written and spoken English and customer-facing communication skills.
  • Preferred experience includes network or storage architecture, high-performance computing or GPU infrastructure, GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand, NVLink, DCGM health checking, GPU driver or firmware lifecycle, NVIDIA AI Enterprise, upstream open source contributions, conference presentations, Rancher, OpenShift, or VMware.

Benefits

  • Professional development and training.
  • Attendance at conferences and working groups.
  • Company outings, happy hours, hackathons, and tech talks.
  • Competitive compensation package with a strong benefits plan.
  • Opportunity to work with open source cloud infrastructure technologies and Fortune 500 and Global 2000 customers.
  • International travel may be required, up to 25%.

Tech Stack

AWSGoJavaScriptKubernetesLinuxOpenShiftOpenStackPythonRancher

Categories

DevOpsSite Reliability
Mirantis

About Mirantis

501-1,000 employees
Contact me