
Site Reliability Engineer
MiQ Digital2 months ago
Bengaluru, IndiaSenior
Responsibilities
- Own the health and reliability of MiQ’s Sigma enterprise platform.
- Design and implement monitoring, alerting, synthetic checks, dashboards, runbooks, SLIs, and SLOs across services and infrastructure.
- Operate Grafana-based observability covering metrics, logs, traces, and synthetic monitoring.
- Partner with product engineering teams to investigate and resolve application, performance, and infrastructure issues.
- Improve release safety through release pipeline optimization, progressive delivery, health gates, deployment analysis, and feature-flag rollbacks.
- Perform load testing, capacity analysis, latency and resource profiling, and other performance engineering activities.
- Automate operational toil with scripting and infrastructure-as-code.
- Apply AIOps capabilities to improve signal quality, detection, diagnosis, and operational efficiency.
- Support incident triage, escalation, communication, post-mortems, and on-call process improvements.
- Grow into ownership of the platform observability charter and establish reliability standards across engineering.
Requirements
- 4–8 years of experience in Site Reliability Engineering, DevOps, or platform/production engineering roles supporting customer-facing systems.
- Hands-on experience with Grafana, Prometheus, Datadog, and log or trace aggregation tools such as Loki, Tempo, or OpenTelemetry.
- Experience creating synthetic API and browser monitoring checks for critical user journeys.
- Deep knowledge of SRE practices including SLIs, SLOs, error budgets, golden signals, alert tuning, noise reduction, and blameless post-incident reviews.
- Experience operating Kubernetes workloads, ideally on EKS, and AWS infrastructure across application, container, and infrastructure layers.
- Strong scripting and automation skills in Python, Bash, or Go.
- Exposure to Terraform, CI/CD pipelines, and infrastructure-as-code practices.
- Experience with load or stress testing tools such as k6, JMeter, or Locust, plus capacity planning and performance profiling.
- Familiarity with AIOps capabilities and their application to SRE maturity.
- Experience with incident management, on-call processes, triage, escalation, communication, and post-mortems.
- Strong written, verbal, communication, collaboration, ownership, and problem-solving abilities.
Benefits
- Hybrid work environment.
- New-hire orientation with job-specific onboarding and training.
- Internal and global mobility opportunities.
- Competitive healthcare benefits.
- Bonus and performance incentives.
- Generous annual PTO and paid parental leave, plus two additional paid days for holidays, cultural events, or inclusion initiatives.
- Employee resource groups supporting connection, action, and communities.
Tech Stack
Categories
Site Reliability