SonicWall

Software Dev Senior Engineer – SRE & Cloud Reliability

SonicWall
Apply
9 days ago
Pune, IndiaSenior

Responsibilities

  • Define and refine SLIs, SLOs, and SLAs for critical services with product and engineering leadership.
  • Own error budget tracking and policies and drive reliability investments when budgets are at risk.
  • Design and implement observability platforms using metrics, structured logging, and distributed tracing.
  • Automate repetitive operational work and lead measurable toil-reduction initiatives.
  • Design and execute chaos engineering programs to identify reliability weaknesses proactively.
  • Facilitate blameless postmortems and track corrective actions to completion.
  • Build incident-response processes, runbooks, escalation paths, and healthy on-call rotations.
  • Design, build, and support resilient, observable, secure, and cost-optimized cloud infrastructure.
  • Automate production deployments on AWS and improve CI/CD pipelines with reliability and quality gates.
  • Develop fault-tolerant management solutions across multiple cloud platforms and data centers.
  • Conduct production-readiness reviews and launch-checklist activities for new services and features.
  • Champion SRE practices, mentor engineers, participate in roadmap planning, and evaluate technologies.

Requirements

  • At least 8 years of experience with scalable, distributed-systems architecture.
  • At least 3 years of hands-on Site Reliability Engineering experience, including ownership of SLOs and error budgets.
  • At least 4 years of cloud-platform experience, including AWS.
  • At least 4 years of infrastructure-as-code experience with Terraform and AWS CDK.
  • At least 5 years of scripting experience using Python, Shell, or a similar language.
  • At least 3 years of containerization experience, including Docker.
  • At least 4 years of orchestration experience, including Kubernetes.
  • Experience designing and operating observability stacks such as Prometheus, Grafana, Datadog, OpenTelemetry, and Jaeger.
  • Experience with incident-management platforms and on-call tooling such as PagerDuty and OpsGenie.
  • Experience implementing automated service deployments covering networking, security, reliability, management, reporting, and configuration management.
  • Experience with chaos-engineering principles and tools such as Chaos Monkey, LitmusChaos, and Gremlin.
  • Experience managing PostgreSQL, Redis, DynamoDB, and MongoDB databases.
  • Strong understanding of deployment automation, production-readiness reviews, networking, HTTP, and distributed-system latency diagnosis.
  • Experience using Git in a team environment and working with distributed tracing frameworks such as Jaeger, Zipkin, and AWS X-Ray.
  • Preferred experience includes Google SRE principles, AIOps or ML-based anomaly detection, capacity planning, load testing, performance benchmarking, CI/CD reliability gates, and customer-facing production issue resolution.
  • CS degree or equivalent experience.

Benefits

  • Hybrid work arrangement in Pune.
  • Equal opportunity employment and commitment to a diverse workplace.

Tech Stack

Amazon DynamoDBAWSDatadogDockerGitGrafanaKubernetesMongoDBPostgreSQLPrometheusPythonRedisTerraform

Categories

DevOpsSite Reliability
SonicWall

About SonicWall

1,001-5,000 employees
Contact me