Mirantis

Senior Site Reliability Engineer

Mirantis
Apply
10 hours ago
Remote, India or Hyderābād, IndiaSenior

Responsibilities

  • Own Kubernetes environments across development, CI, pre-production, and a customer-facing production region, including multi-cluster and multi-region topologies.
  • Operate the production region against SLOs, including capacity planning, upgrades, patching, backup and restore, disaster recovery drills, and on-call participation.
  • Lead incident detection, mitigation, customer-impact assessment, root-cause analysis, and blameless postmortems.
  • Build and maintain Helm charts, umbrella releases, versioning practices, values management, and upgrade paths.
  • Own CI/CD workflows for builds, testing, image publishing, chart packaging, releases, hotfixes, and backports.
  • Automate environment bootstrap and full-stack seeding for the control plane, identity, gateway, database, and workflow engine.
  • Operate and troubleshoot PostgreSQL, Temporal, Keycloak, API gateways, message brokers, and observability components.
  • Build metrics, dashboards, alerting, logging, audit access, diagnostics, and SLO-oriented observability.
  • Advise product and regional operations teams on deployment topology, GPU scheduling, RBAC, networking, failure modes, upgrades, and migrations.
  • Enforce least-privilege access, secret handling, certificate and TLS lifecycle management, image and dependency scanning, tenant isolation, and compliance evidence collection.
  • Drive infrastructure as code, repeatable environments, documentation, runbooks, and operational best practices.
  • Mentor engineers on Kubernetes and operational practices and improve team standards through reviews and documentation.

Requirements

  • 10+ years of experience in DevOps, SRE, platform, or infrastructure engineering, including production ownership of customer-facing Kubernetes environments.
  • Expert-level Kubernetes knowledge covering workloads, networking, storage, RBAC, resource management, CRDs, operators, and cluster upgrades.
  • Proven incident-response experience involving on-call rotations, escalation paths, postmortems, and corrective actions.
  • Strong CI/CD engineering experience, including pipelines as code, reproducible builds, artifact management, and release management.
  • Strong scripting and automation skills with enough Go familiarity to read service code, trace failures, and file precise bugs.
  • Experience operating stateful services such as relational databases, identity providers, gateways, and message brokers in Kubernetes, including backup, restore, and upgrades.
  • Experience consulting with engineering teams through runbooks, design feedback, and incident write-ups across global time zones; strong written English.
  • Experience with Kubernetes at scale, Cluster API, controllers/operators, Docker, and Helm chart authoring and lifecycle management.
  • Experience with GitHub Actions or equivalent CI/CD tooling, container registries, versioned releases, and backport workflows.
  • Experience with Terraform, Ansible, or equivalent infrastructure-as-code tools and GitOps tooling such as Argo CD or Flux.
  • Experience operating Keycloak and API gateways, including routing, plugins, TLS, and rate limiting.
  • Experience with PostgreSQL operations and migrations and streaming or message-broker platforms such as Kafka.
  • Experience with Prometheus, Grafana, centralized logging, and alerting tied to SLOs.
  • Experience with AWS networking, IAM, load balancing, and managed Kubernetes.
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or 10 years of related experience.
  • Nice-to-have experience includes k0s or k0rdent, GPU infrastructure, Temporal, Metal³, BareMetalHost, OpenStack, multi-region systems, service mesh, cross-cluster networking, Python, pytest, performance testing, OPA, Kyverno, OpenTelemetry, distributed tracing, compliance audits, and CNCF contributions.

Benefits

  • Professional development and training.
  • Opportunities to attend conferences and working groups.
  • Company outings, happy hours, hackathons, and tech talks.
  • Work with an established Silicon Valley cloud infrastructure leader and passionate colleagues serving Fortune 500 and Global 2000 customers.
  • Participation in open-source innovation and a collaborative, high-energy work environment.
  • Competitive compensation package with a strong benefits plan.

Tech Stack

AnsibleApache KafkaArgo CDAWSDockerGitHub ActionsGoGrafanaHelmKubernetesOpenStackPostgreSQLPrometheuspytestPythonTerraform

Categories

Site Reliability
Mirantis

About Mirantis

501-1,000 employees

Mirantis builds Kubernetes-native infrastructure software and services for enterprises to run cloud, container, and AI workloads across on-premises, public cloud, and edge environments. It sells subscriptions, support, and managed services around products like Mirantis Kubernetes Engine and the Lens Kubernetes platform; the company acquired Docker Enterprise in 2019. Founded in 1999 and headquartered in Campbell, California, Mirantis is privately held and an IREN company serving global enterprises.

Contact me