
Site Reliability Engineer
HighRadius15 days ago
Hyderābād, IndiaSenior
Responsibilities
- Design, build, maintain, and optimize resilient cloud infrastructure for scalable and reliable applications.
- Establish and manage monitoring and alerting systems and use SLI, SLO, SLA, and error-budget practices to improve service reliability.
- Automate infrastructure provisioning, system configuration, repetitive operational tasks, and workflows; develop standard operating procedures.
- Participate in major incident response, perform root cause analyses, and implement permanent production solutions.
- Collaborate with software engineers and cross-functional stakeholders on cloud integrations, requirements, proof of concepts, resilience, and common tooling.
- Promote and implement tools, components, and processes that improve resilience and reduce manual operational effort.
Requirements
- At least 7 years of industry experience.
- Expertise in at least one of AWS, Azure, or GCP, plus extensive experience with Docker and Kubernetes.
- Proficiency with shell scripting and Python; experience with Ansible or Puppet and Terraform or CloudFormation.
- Experience with Prometheus, Grafana, or the ELK stack, including hands-on OpenTelemetry implementation.
- Strong understanding of SLI, SLO, SLA, and error budgeting.
- Experience with on-premises hosting and virtualization using VMware, Hyper-V, or KVM.
- Understanding of storage technologies and protocols including NAS, SAN, EFS, NFS, FTP, SFTP, SMTP, NTP, DNS, and DHCP.
- Experience with networking, firewall technologies, Linux internals, RHEL, CentOS, Rocky Linux, and Windows operating systems.
- Bonus experience with Git, Jenkins, Rundeck, ArgoCD, Crossplane, MySQL, Hadoop, disaster recovery, business continuity, cloud cost optimization, cloud security, and ITIL processes.
- Strong communication, collaboration, adaptability, proactive problem-solving, and reliability advocacy skills.
Tech Stack
AnsibleApache HadoopAWSAzureDockerGitGoogle Cloud PlatformGrafanaJenkinsKubernetesLinuxMySQLPrometheusPuppetPythonTerraformWindows
Categories
Site Reliability
About HighRadius
HighRadius builds an AI-driven finance platform for enterprises, spanning order-to-cash, accounts receivable, cash application, deductions, collections, credit, accounts payable, close/reconciliation, and treasury. It sells its software as a cloud SaaS offering with implementation and support services. Founded in 2006 and headquartered in Houston, its customers include 3M, Unilever, Red Bull, and Lufthansa.