4 days ago
Pune, IndiaMid Level
Responsibilities
- Monitor platform health and improve reliability, performance, availability, and operational efficiency.
- Perform incident triage, troubleshooting, stakeholder communication, and post-incident reviews.
- Design and implement scalable, resilient solutions with engineering, infrastructure, and business teams.
- Build deployment and observability capabilities, including dashboards, alerts, and service health monitoring.
- Analyze logs, metrics, and distributed traces to identify bottlenecks and reliability issues.
- Support Kafka-based event-driven platforms, enterprise messaging, and integrations.
- Apply capacity planning and availability management practices to improve resilience.
- Develop automation to reduce operational toil and manual intervention.
- Participate in production support, problem management, release management, and change management.
Requirements
- Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related discipline.
- At least 3 years of experience in SRE, Platform Reliability Engineering, DevOps, production support, or application support.
- Strong programming or scripting experience in Python, Go, or Java.
- Knowledge of software engineering principles, data structures, algorithms, and system design.
- Experience with Linux/Unix and Windows Server environments.
- Experience with monitoring, observability, reliability concepts, metrics, logs, traces, SLIs, SLOs, and alerting.
- Understanding of event-driven architectures and enterprise messaging platforms such as Kafka and MQ.
- Experience troubleshooting distributed production systems, APIs, middleware components, and message flows.
- Understanding of incident management, problem management, root cause analysis, and operational support processes.
- Familiarity with source control, CI/CD pipelines, Infrastructure as Code, and DevOps practices.
- Strong verbal and written communication skills and the ability to work with technical and business stakeholders.
- Preferred experience with Grafana, Prometheus, OpenTelemetry, Loki, Git, Ansible, Docker, Kubernetes, Kafka, Redis, MQ, and AWS.
Tech Stack
Categories
Site Reliability
About Jefferies
Jefferies is a global investment banking and capital markets firm serving corporations, financial institutions, governments, and individuals with M&A advisory, underwriting, sales and trading, research, and wealth management. Founded in 1962 and headquartered in New York, it operates from 40+ offices worldwide. The firm earns fees and trading revenues and is owned by publicly traded Jefferies Financial Group (NYSE: JEF).
