2 hours ago
Hyderābād, IndiaStaff+
Responsibilities
- Ensure high availability, scalability, resilience, and performance of APIs, .NET applications, telemetry ingestion pipelines, and on-premise infrastructure.
- Monitor production health, availability, error rates, resource saturation, latency, throughput, and end-to-end performance.
- Define and operationalize SLIs, SLOs, SLAs, error budgets, availability models, and capacity-planning practices.
- Build software and automation for infrastructure management, deployments, application operations, and reduction of manual toil.
- Own incident detection, triage, mitigation, communication, root-cause analysis, and post-mortems.
- Design and maintain monitoring, logging, alerting, distributed tracing, and observability frameworks.
- Develop CI/CD pipelines and integrate reliability practices into the software development lifecycle.
- Manage Rancher, OpenShift, Kubernetes, virtualization, storage, networking, and security controls in on-premise environments.
- Provision and maintain infrastructure using infrastructure-as-code, automation scripts, and configuration management tools.
- Optimize backend services, telemetry systems, and high-load ingestion pipelines for capacity, performance, and fleet-operation load spikes.
- Support real-time vehicle telemetry, IoT, GPS/GNSS, connectivity, message delivery, offline synchronization, and edge-constrained data flows.
- Maintain architecture documentation, operational processes, incident playbooks, and system runbooks.
Requirements
- Bachelor’s degree in Computer Science or equivalent.
- 8–12 years of relevant experience in DevOps or Site Reliability Engineering.
- Hands-on experience operating production systems, CI/CD pipelines, and distributed application platforms.
- Strong expertise with on-premise infrastructure, container orchestration, networking, storage, IP routing, firewalls, and security controls.
- Deep observability and monitoring experience covering distributed services, .NET and Java applications, and telemetry pipelines.
- Advanced reliability engineering experience defining SLI/SLO/SLA practices, error budgets, availability models, and capacity and performance plans.
- Strong automation and CI/CD experience with GitHub, Octopus, Jenkins or Azure DevOps, Terraform, Helm, Kustomize, PowerShell, Bash, and Python.
- Production operations experience with incident management, system-health monitoring, performance analysis, scalability improvements, and uptime SLAs.
- Backend performance and systems engineering experience, including .NET profiling, SQL Server and Redis tuning, and high-load telemetry ingestion.
- Experience supporting real-time data flows, IoT or telemetry systems, GPS/GNSS streams, vehicle connectivity, offline or edge constraints, and mobility-driven scaling.
- Knowledge of secrets management, least-privilege access, vulnerability scanning, data protection, security, and regulatory compliance.
- Experience with AWS services such as EKS, EC2, RDS, S3, VPC, and IAM is preferred, especially in hybrid on-premise and cloud environments.
Tech Stack
Categories
DevOpsSite Reliability
