
Associate Site Reliability Engineer
Amtech Software6 days ago
Bengaluru, IndiaEntry Level
Responsibilities
- Monitor and tune production monitoring, logging, and alerting using Open Observe and CloudWatch.
- Track service-level indicators and objectives and help maintain reliability dashboards and reports.
- Assist with incident root cause analysis and document findings and follow-up actions.
- Automate routine operational tasks with Python or Bash and contribute to Terraform modules under senior review.
- Support containerized and serverless workloads on ECS Fargate, EKS, and Lambda.
- Support GitHub Actions CI/CD workflows and assist with environment provisioning.
- Create and maintain runbooks and operational playbooks.
- Participate in the on-call rotation, respond to PagerDuty alerts, and execute documented recovery procedures.
- Recommend improvements to reduce mean time to detection and mean time to recovery.
- Apply IAM, secrets management, least-privilege practices, and applicable security and compliance policies.
- Use and verify AI-assisted engineering tools safely for scripting, documentation, and troubleshooting.
- Learn through paired work, design reviews, post-incident reviews, and internal knowledge-base contributions.
Requirements
- 1–2 years of experience in SRE, DevOps, systems, or cloud engineering; internships and substantial personal or academic projects count.
- Ability to script in Python or Bash and reason methodically through technical problems.
- Foundational knowledge of AWS core services including EC2, S3, IAM, and VPC, or equivalent cloud fundamentals.
- Familiarity with Linux administration, basic networking including DNS, HTTP, and TCP/IP, Windows Server, and VMware Infrastructure.
- Exposure to Git-based workflows such as GitHub or similar tools.
- Strong analytical and troubleshooting abilities and willingness to learn from production systems.
- Safe and effective use of AI assistants, including verifying output before use.
- Bachelor’s degree in Computer Science, Engineering, or a related discipline, or equivalent demonstrated skills.
- Preferred: AWS Certified Cloud Practitioner or Solutions Architect Associate certification.
- Preferred: exposure to Terraform, Docker, or Kubernetes.
- Preferred: understanding of SLO/SLI concepts and incident management workflows.
- Preferred: experience with CloudWatch, Prometheus, Grafana, or OpenTelemetry-based monitoring stacks.
Benefits
- Structured mentorship, paired work, design reviews, and opportunities to progress through the engineering career ladder.
- Access to continuous learning, leadership development, and cross-portfolio career opportunities through Vista’s global network.
- Collaborative, transparent, people-first culture with exposure to a growing enterprise software organization and Vista’s global ecosystem.
Categories
Site Reliability