10 hours ago
Remote, Poland or Kraków, PolandSenior
Responsibilities
- Develop and maintain automated tools and scripts for system reliability, deployment, and incident response.
- Improve monitoring, error detection, remediation, performance, and reliability of virtualization infrastructure.
- Participate in on-call rotations and guide restoration of service-impacting issues.
- Write automation and tooling to reduce operational toil and improve deployment safety.
- Contribute to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure.
- Mentor other engineers and promote reliability and operational excellence across teams.
Requirements
- Expert-level experience in Linux/Unix administration, DevOps, or SRE roles involving large-scale distributed systems.
- Expertise with Kubernetes and large-scale containerization systems.
- Programming experience in Python or Golang.
- Experience with Terraform, SaltStack, or Ansible for configuration management.
- Experience defining SLOs and using observability tools such as Prometheus and Grafana, along with distributed tracing.
- Experience architecting software and infrastructure at scale and developing automation and monitoring.
Benefits
- Career development opportunities through programs such as GROW and Mentoring, internal events such as APEX Expo, and tools such as LinkedIn Learning.
- A 15-minute exploratory call with a recruiter is available for candidates seeking more information.
About Akamai
Akamai builds content delivery, cloud security, and edge compute services for enterprises that run consumer and business web, API, and video workloads. The public company (NASDAQ: AKAM), founded in 1998 and headquartered in Cambridge, MA, operates a global CDN and security platform and sells via subscriptions and usage-based services. It expanded into cloud infrastructure by acquiring Linode in 2022, offering IaaS and managed Kubernetes.
