
Site Reliability Engineer, Cloud Platform
Qualys, Inc.2 hours ago
Pune, IndiaMid Level
Responsibilities
- Co-develop and operate cloud platform services across design, deployment, production operation, and improvement.
- Improve platform reliability, performance, availability, scalability, and effectiveness through measurement and automated changes.
- Perform system design, capacity planning, deployment automation, production monitoring and alerting, and testing or verification.
- Develop tools and automate large-scale provisioning and deployment of cloud platform technologies.
- Monitor latency, performance, system health, and anomalous behavior, and identify root causes of incidents.
- Participate in on-call rotations, lead incident response, and write detailed postmortem analyses.
- Drive improvements in capacity planning, configuration management, scaling, performance tuning, monitoring, alerting, and root-cause analysis.
Requirements
- 4+ years of relevant experience running distributed systems at scale in production.
- Expertise in at least one of Java, Python, or Go.
- Proficiency writing Bash scripts.
- Understanding of SQL and NoSQL systems, systems programming, network stacks, file systems, and operating-system services.
- Understanding of firewalls, load balancers, DNS, NAT, TLS/SSL, and VLANs.
- Ability to identify performance bottlenecks and anomalous system behavior and determine incident root causes.
- Knowledge of JVM concepts including garbage collection, heap, stack, profiling, and class loading.
- Knowledge of security, performance, high availability, and disaster recovery best practices.
- Proven experience handling production issues, planning escalation procedures, conducting postmortems, performing impact analysis, and completing risk assessments.
- Ability to drive results and set priorities independently.
- BS or MS degree in Computer Science, Applied Math, or a related field.