
Senior DevOps Engineer
Qualys, Inc.1 hour ago
Pune, IndiaSenior
Responsibilities
- Design, build, operate, and improve scalable observability platforms using ClickHouse, HyperDX, OpenTelemetry, Kubernetes, Prometheus, Grafana, Alertmanager, Fluent Bit, and Filebeat.
- Architect and optimize ClickHouse schemas, partitioning, ordering keys, TTLs, materialized views, retention policies, storage efficiency, and queries for high-volume observability workloads.
- Operate HyperDX for log search, distributed tracing, dashboards, service analysis, telemetry correlation, troubleshooting, and root-cause analysis.
- Build and maintain OpenTelemetry Collector pipelines for collecting, processing, enriching, filtering, sampling, and routing telemetry.
- Design resilient telemetry pipelines with batching, queuing, retries, backpressure handling, sampling, rate limiting, and cardinality controls.
- Deploy and operate observability infrastructure on Kubernetes with emphasis on scalability, availability, capacity planning, resiliency, and operational safety.
- Automate infrastructure deployment, configuration management, platform upgrades, application onboarding, and recurring operational work.
- Build and maintain Jenkins CI/CD workflows and use Terraform, Ansible, HashiCorp Consul, and HashiCorp Vault for automation and integrations.
- Troubleshoot production issues across applications, infrastructure, telemetry pipelines, storage systems, and ClickHouse queries, and drive corrective actions to completion.
- Define standards for telemetry instrumentation, log quality, metric hygiene, trace propagation, dashboards, alert quality, and production readiness.
- Participate in incident response, post-incident reviews, capacity planning, operational reviews, and remediation tracking.
- Mentor junior engineers, review technical designs and automation changes, document operational procedures, and lead complex incident investigations.
Requirements
- At least 6 years of experience in DevOps, SRE, Platform Engineering, Infrastructure Engineering, or Observability Engineering roles.
- Strong production experience with ClickHouse architecture, administration, schema design, performance tuning, query optimization, retention management, and operations at scale.
- Hands-on experience deploying and operating HyperDX with ClickHouse for log search, distributed tracing, dashboards, service analysis, and troubleshooting.
- Strong experience with OpenTelemetry and OpenTelemetry Collector, including receiver, processor, exporter, sampling, enrichment, batching, and routing configurations.
- Strong working knowledge of Fluent Bit, Filebeat, Prometheus, Alertmanager, Grafana, ClickHouse, HyperDX, OpenTelemetry, Jenkins, Ansible, Terraform, HashiCorp Consul, HashiCorp Vault, and Kubernetes.
- Understanding of Linux, networking, microservices, REST, gRPC, distributed systems, high-availability architectures, and production operations.
- Experience operating and troubleshooting high-volume production systems with focus on reliability, scalability, capacity, and performance.
- Ability to debug issues across application telemetry, Kubernetes infrastructure, data ingestion pipelines, storage systems, and query performance.
- Strong scripting and automation skills and the ability to improve repeatability, reliability, and operational efficiency.
- Ability to own production systems, drive issues to closure, communicate during incidents, and collaborate with engineering and operations teams.
- Preferred development experience with Java and Python.
- Preferred experience developing automation, internal tools, APIs, platform services, or self-service onboarding workflows.
- Preferred experience with Kafka or other high-throughput messaging and streaming platforms.
- Preferred understanding of APM, distributed tracing, OpenTelemetry instrumentation, context propagation, service maps, service dependency analysis, SLOs, and SLIs.
- Preferred experience with AWS, Azure, GCP, or OCI and with observability solutions for large-scale Kubernetes, microservices, or distributed application environments.
Tech Stack
AnsibleApache KafkaAWSAzureClickHouseGoogle Cloud PlatformGrafanagRPCJavaJenkinsKubernetesLinuxPrometheusPythonTerraform