
Senior Site Reliability Engineer
Truecaller18 days ago
Bengaluru, IndiaSenior
Responsibilities
- Design, manage, and maintain infrastructure services throughout their lifecycle.
- Build tooling to provision and scale infrastructure resources.
- Improve the scalability, performance, reliability, availability, security, and evolution of infrastructure components.
- Develop instrumentation, tooling, and alerting systems to monitor availability and escalate abnormalities.
- Support monitoring and capacity planning with application development teams in alignment with business goals.
- Respond to incidents, perform root cause analysis, and implement improvements to system reliability.
- Deploy, scale, and monitor Kubernetes clusters and automate configuration, deployment, and service management activities.
Requirements
- Extensive Linux system administration knowledge, preferably with high-throughput and low-latency systems.
- Strong hands-on experience with GCP services or transferable AWS/Azure skills, including networking, IAM, compute, storage, and Kubernetes.
- Extensive knowledge of Docker and Kubernetes, including cluster deployment, scaling, and monitoring.
- Excellent understanding of distributed system design across process and site boundaries.
- Hands-on experience with service orchestration, management, deployment, configuration management, and automation.
- Strong understanding of process isolation and containerization concepts.
- Experience with monitoring and observability tools such as Prometheus, Grafana, Stackdriver, Datadog, and New Relic.
- Practical incident management experience, including incident response, root cause analysis, and reliability improvements.
- Knowledge of cloud security, secrets management, and compliance basics.
- Understanding of software development lifecycle, versioning, building, testing, staging, deployment, and continuous delivery processes.
- Preferred experience developing Kubernetes operators.
- Preferred experience deploying and scaling Apache Cassandra, ScyllaDB, MySQL, PostgreSQL, Redis, or Memcached.
- Go programming experience or willingness to learn Go.
Benefits
- Learning resources, leadership programs, mentoring, hands-on work, internal mobility, and a transparent progression path.
- Learning and development allowance, voluntary provident fund and/or national pension scheme tax-saving options, and creche allowance.
- Choice of computer and phone within the company budget.
- Office-first work model with some flexibility and in-person collaboration facilities.
- Breakfast, lunch, quiet spaces, team activities, movie nights, tech meetups, and cultural events.
- Five quarterly Lab Days for exploration, prototyping, and building new ideas.
- Applications must be submitted in English, and the recruitment process includes a background check.
Tech Stack
Apache CassandraAWSAzureDatadogDockerGoGoogle Cloud PlatformGrafanaKubernetesLinuxMySQLPostgreSQLPrometheusRedis
Categories
Site Reliability