about 5 hours ago
Responsibilities
- Lead the design and implementation of reliable and scalable production platforms.
- Collaborate with cross-functional teams to maintain resilient infrastructure.
- Provide technical leadership and mentorship to engineers.
- Participate in a 24x7 on-call rotation for critical services.
- Drive standardization, automation, and documentation efforts.
- Contribute to the full lifecycle of platform and service delivery.
Requirements
- 5+ years of experience in DevOps, SRE, or platform engineering roles.
- Strong Kubernetes experience at scale.
- Hands-on experience with infrastructure as code tools like Terraform or Ansible.
- Strong programming skills in at least one object-oriented language.
- Understanding of security principles across infrastructure and services.
- Significant experience in at least one major cloud platform.
- Experience with monitoring and observability tools like Prometheus.
- Solid understanding of networking fundamentals and distributed systems.
- Strong Linux and/or Windows systems administration experience.
- Experience with CI/CD pipelines and secure SDLC practices.
- Good understanding of SRE concepts such as SLIs and SLOs.
Tech Stack
AnsibleApache KafkaAWSElasticsearchGoogle Cloud PlatformGrafanaKubernetesLinuxMySQLPostgreSQLPrometheusPuppetRedisTerraformWindows
