
[MLA] Senior Site Reliability Engineer (SRE) – Kubernetes
Software Mind22 days ago
Remote, Poland or Kraków, PolandSenior
Responsibilities
- Deploy, operate, and maintain reliable production services running on Kubernetes.
- Monitor service health and investigate incidents across distributed applications.
- Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements.
- Troubleshoot application runtime, networking, and service-to-service issues across Node.js and JVM-based systems.
- Support CI/CD, GitOps deployments, observability, production monitoring, alerting, and dashboards.
- Work within a client-directed backlog and established priorities.
Requirements
- At least 5 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering, or a closely related role.
- Strong recent hands-on experience supporting Kubernetes-based production services, with 3+ years of production Kubernetes experience strongly preferred.
- Experience with Kubernetes deployment, scaling, rollouts and rollbacks, resource tuning, and service-to-service troubleshooting.
- Strong production incident response experience, including on-call support, runbooks, postmortems, and paging hygiene.
- Experience with Splunk, Prometheus, and Grafana, including building alert rules and dashboards.
- Experience with CI/CD and infrastructure-as-code for containerized deployments, including Helm and GitOps tools such as ArgoCD or Flux.
- Strong Linux and networking fundamentals, including DNS, load balancing, TCP, HTTP, HTTP/2, and Kubernetes networking.
- Production troubleshooting experience across Node.js and JVM/Java services, with strong depth in at least one runtime environment.
- Experience with service-to-service authentication, mTLS, certificate rotation, certificate format conversion, and JWT-based authentication.
- Very good spoken and written English.
- Preferred experience includes Lit or Web Components, server-side rendering or isomorphic runtimes, canary rollouts, multi-version production operations, distributed tracing, request-context correlation, KEDA, and enterprise platform integration layers.
Benefits
- Flexible employment and remote work.
- International projects with leading global clients and opportunities for international business trips.
- Language classes and internal and external training.
- Private healthcare and insurance.
- Multisport card and well-being initiatives.
- Non-corporate work environment.
Tech Stack
Categories
DevOpsSite Reliability