
Site Reliability Engineer, Cloud Platform
Qualys, Inc.1 day ago
Pune, IndiaEntry Level / Mid Level
Responsibilities
- Monitor production applications, infrastructure, Kubernetes environments, and distributed systems to maintain service health and reliability.
- Participate in incident response, troubleshooting, escalation, root-cause analysis, post-incident reviews, and on-call rotations.
- Support Apache Spark workloads, Kafka pipelines, cloud-based data platforms, and distributed data-processing systems.
- Assist with Kubernetes deployments, containerized application troubleshooting, cloud operations, configuration, and operational workflows.
- Create and maintain dashboards, alerts, monitoring configurations, runbooks, standard operating procedures, and troubleshooting guides.
- Develop Python, Bash, Go, or similar automation and tooling for monitoring, log analysis, deployments, and operational tasks.
- Support CI/CD, Infrastructure as Code, automated deployment workflows, capacity planning, and self-service capabilities.
- Collaborate with Engineering, DevOps, Infrastructure, Security, and Data Platform teams to improve reliability, observability, scalability, and operational efficiency.
Requirements
- 0–3 years of experience in SRE, DevOps, Cloud Engineering, Systems Engineering, or a related technical role.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- Basic understanding of Linux/Unix systems, processes, memory, networking, and system troubleshooting.
- Fundamental understanding of distributed systems and cloud-native technologies, with experience or familiarity with Kubernetes and containerized applications.
- Familiarity with Apache Spark, Kafka, or similar distributed data-processing platforms.
- Programming and scripting fundamentals in Python, Bash, Go, Java, Scala, or similar languages, with the ability to write automation scripts.
- Basic SQL knowledge and strong problem-solving, troubleshooting, analytical, communication, and collaboration skills.
- Familiarity with monitoring and observability tools such as Prometheus and Grafana.
- Basic understanding of cloud platforms such as AWS, OCI, or Google Cloud Platform.
- Familiarity with Delta Lake, S3, distributed data platforms, CI/CD, automated deployments, Terraform, CloudFormation, or other Infrastructure as Code tools is preferred.
- Experience with scripting automation, internal tooling, large-scale or distributed systems, incident management, on-call processes, or postmortems is preferred.
- Familiarity with Jenkins, GitHub, Bitbucket, ELK, Splunk, or AppDynamics and cloud certifications are preferred.
Benefits
- Full-time role with hybrid or remote flexibility depending on company policy.
- On-call participation may be required.
- Opportunity to work with cloud-native, distributed, and large-scale production systems.
- Cross-functional collaboration with Engineering, DevOps, Infrastructure, Security, and Data Platform teams.
Tech Stack
Apache KafkaApache SparkAWSBashGoGoogle Cloud PlatformGrafanaJavaJenkinsKubernetesLinuxOracle CloudPrometheusPythonScalaSplunkSQLTerraform
Categories
DevOpsSite Reliability
About Qualys, Inc.
Qualys builds a cloud-based security and compliance platform used by enterprises to automate vulnerability management, IT asset visibility, web application scanning, and policy/PCI compliance across on‑prem and cloud environments. Delivered as subscription SaaS, its Enterprise TruRisk Platform uses a single agent to provide continuous security telemetry for endpoints, servers, containers, and clouds. Founded in 1999 and headquartered in Foster City, California, Qualys is a public company trading on NASDAQ (QLYS).