10 hours ago
Remote, India or Hyderābād, IndiaSenior
Responsibilities
- Own Kubernetes environments across development, CI, pre-production, and a customer-facing production region, including multi-cluster and multi-region topologies.
- Operate the production region against SLOs, including capacity planning, upgrades, patching, backup and restore, disaster recovery drills, and on-call participation.
- Lead incident detection, mitigation, customer-impact assessment, root-cause analysis, and blameless postmortems.
- Build and maintain Helm charts, umbrella releases, versioning practices, values management, and upgrade paths.
- Own CI/CD workflows for builds, testing, image publishing, chart packaging, releases, hotfixes, and backports.
- Automate environment bootstrap and full-stack seeding for the control plane, identity, gateway, database, and workflow engine.
- Operate and troubleshoot PostgreSQL, Temporal, Keycloak, API gateways, message brokers, and observability components.
- Build metrics, dashboards, alerting, logging, audit access, diagnostics, and SLO-oriented observability.
- Advise product and regional operations teams on deployment topology, GPU scheduling, RBAC, networking, failure modes, upgrades, and migrations.
- Enforce least-privilege access, secret handling, certificate and TLS lifecycle management, image and dependency scanning, tenant isolation, and compliance evidence collection.
- Drive infrastructure as code, repeatable environments, documentation, runbooks, and operational best practices.
- Mentor engineers on Kubernetes and operational practices and improve team standards through reviews and documentation.
Requirements
- 10+ years of experience in DevOps, SRE, platform, or infrastructure engineering, including production ownership of customer-facing Kubernetes environments.
- Expert-level Kubernetes knowledge covering workloads, networking, storage, RBAC, resource management, CRDs, operators, and cluster upgrades.
- Proven incident-response experience involving on-call rotations, escalation paths, postmortems, and corrective actions.
- Strong CI/CD engineering experience, including pipelines as code, reproducible builds, artifact management, and release management.
- Strong scripting and automation skills with enough Go familiarity to read service code, trace failures, and file precise bugs.
- Experience operating stateful services such as relational databases, identity providers, gateways, and message brokers in Kubernetes, including backup, restore, and upgrades.
- Experience consulting with engineering teams through runbooks, design feedback, and incident write-ups across global time zones; strong written English.
- Experience with Kubernetes at scale, Cluster API, controllers/operators, Docker, and Helm chart authoring and lifecycle management.
- Experience with GitHub Actions or equivalent CI/CD tooling, container registries, versioned releases, and backport workflows.
- Experience with Terraform, Ansible, or equivalent infrastructure-as-code tools and GitOps tooling such as Argo CD or Flux.
- Experience operating Keycloak and API gateways, including routing, plugins, TLS, and rate limiting.
- Experience with PostgreSQL operations and migrations and streaming or message-broker platforms such as Kafka.
- Experience with Prometheus, Grafana, centralized logging, and alerting tied to SLOs.
- Experience with AWS networking, IAM, load balancing, and managed Kubernetes.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or 10 years of related experience.
- Nice-to-have experience includes k0s or k0rdent, GPU infrastructure, Temporal, Metal³, BareMetalHost, OpenStack, multi-region systems, service mesh, cross-cluster networking, Python, pytest, performance testing, OPA, Kyverno, OpenTelemetry, distributed tracing, compliance audits, and CNCF contributions.
Benefits
- Professional development and training.
- Opportunities to attend conferences and working groups.
- Company outings, happy hours, hackathons, and tech talks.
- Work with an established Silicon Valley cloud infrastructure leader and passionate colleagues serving Fortune 500 and Global 2000 customers.
- Participation in open-source innovation and a collaborative, high-energy work environment.
- Competitive compensation package with a strong benefits plan.
Tech Stack
AnsibleApache KafkaArgo CDAWSDockerGitHub ActionsGoGrafanaHelmKubernetesOpenStackPostgreSQLPrometheuspytestPythonTerraform
Categories
Site Reliability
About Mirantis
Mirantis builds Kubernetes-native infrastructure software and services for enterprises to run cloud, container, and AI workloads across on-premises, public cloud, and edge environments. It sells subscriptions, support, and managed services around products like Mirantis Kubernetes Engine and the Lens Kubernetes platform; the company acquired Docker Enterprise in 2019. Founded in 1999 and headquartered in Campbell, California, Mirantis is privately held and an IREN company serving global enterprises.
