1 day ago
Remote, IndiaMid Level
Responsibilities
- Design, implement, and maintain cloud infrastructure on AWS and Google Cloud.
- Ensure the scalability, performance, and reliability of Kubernetes-based distributed database systems.
- Write production-grade Golang code and automation for infrastructure management and system operations.
- Optimize CI/CD pipelines, deployment processes, monitoring, logging, and alerting systems.
- Develop disaster recovery, high-availability, and fault-tolerance strategies.
- Identify bottlenecks and troubleshoot complex issues across networks, operating systems, and cloud infrastructure.
- Participate in on-call rotations and respond to critical production incidents.
- Collaborate with development, cross-functional, and Customer Success teams to improve reliability and resolve customer issues.
Requirements
- Experience in an SRE or DevOps Engineer role in cloud-native environments.
- Advanced knowledge of AWS or Google Cloud, Linux internals, Docker, and Kubernetes at scale.
- Experience with Jenkins or CircleCI, Prometheus, Grafana, the ELK stack, and Git.
- Knowledge of core networking, cloud security best practices, and systematic infrastructure troubleshooting.
- Programming proficiency in Golang or Python.
- Strong communication, autonomy, organization, and remote-team collaboration skills.
- Experience with distributed databases or large-scale data storage systems is preferred.
- Python or Bash scripting, Terraform, GitOps, and strong Golang automation experience are preferred.
Benefits
- Remote position based in India.
- Opportunity to work on AI and data infrastructure and help shape enterprise AI applications.
- Collaboration with experienced engineering, marketing, and product teams.
Tech Stack
AWSBashCircleCIDockerGitGoGoogle CloudGoogle Cloud PlatformGrafanaJenkinsKubernetesLinuxPrometheusPythonTerraform
Categories
DevOpsSite Reliability
