Site Reliability Engineer
LambdaTest21 days ago
Noida, IndiaEntry Level / Mid Level
Responsibilities
- Own and scale production AWS services and cloud infrastructure.
- Manage Kubernetes workloads on EKS and implement Karpenter and KEDA autoscaling.
- Build and optimize CI/CD pipelines using Jenkins, Docker Buildx, ECR caching, ArgoCD, and Helm.
- Own application performance monitoring, production monitoring, debugging, and incident response.
- Write custom code and automation solutions for live production systems.
- Provision infrastructure with Terraform and manage state and concurrent-apply risks.
Requirements
- 1–3 years of hands-on DevOps or cloud infrastructure experience.
- Production experience with AWS, including EKS, SQS, ECR, Route 53, and ALB/NLB.
- Production experience with Docker and Kubernetes.
- Hands-on experience with at least one of New Relic, Sumo Logic, Prometheus, or Grafana.
- Real production incident-management experience with the ability to explain root causes.
- Programming ability in any language and the ability to build custom automation solutions.
- Ability to explain architectural decisions and bridge development and DevOps perspectives.
- Golang or Java backend coding experience is preferred.
- Experience with KEDA, Karpenter, ArgoCD, Helm, Istio, Kafka, or SQS is preferred.
- Awareness of service-layer system design is preferred.
Benefits
- Full-time role located in Noida.
- Direct exposure to production systems at scale and ownership from day one.
- Clear growth path in Platform and SRE Engineering.
Tech Stack
Categories
DevOpsSite Reliability