4 days ago
Pune, IndiaMid Level
Responsibilities
- Design, build, and maintain AI tooling infrastructure running on AWS Kubernetes.
- Monitor platform health and improve system reliability, performance, scalability, and availability.
- Perform incident triage, troubleshooting, stakeholder communication, and post-incident reviews.
- Develop automation to reduce operational toil and manual support activities.
- Build deployment, monitoring, alerting, dashboards, and observability capabilities.
- Use Grafana, Datadog, Prometheus, and OpenTelemetry to monitor service health and analyze platform behavior.
- Analyze logs, metrics, and distributed traces to identify bottlenecks and reliability issues.
- Support Kafka-based enterprise messaging and event-driven architectures.
- Troubleshoot applications, middleware, databases, messaging platforms, infrastructure, and cloud environments.
- Participate in production support, problem management, release management, and change management activities.
- Collaborate with global teams across APAC, EMEA, and the Americas on technology initiatives.
Requirements
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline.
- At least 3 years of experience in Site Reliability Engineering, Platform Reliability Engineering, DevOps, Production Support, or Application Support.
- Strong programming and scripting experience in Python, Go, C#, Java, or C++.
- Understanding of software engineering principles, data structures, algorithms, and system design.
- Experience supporting and troubleshooting distributed applications in production environments.
- Working knowledge of Linux/Unix and Windows Server environments.
- Experience with relational and NoSQL databases, including performance analysis and troubleshooting.
- Hands-on experience with Grafana, Datadog, Prometheus, OpenTelemetry, Loki, and Jaeger.
- Experience creating operational dashboards, alerts, runbooks, and monitoring solutions.
- Understanding of event-driven architectures and enterprise messaging platforms such as Kafka.
- Ability to troubleshoot message flows, APIs, middleware components, and distributed systems.
- Understanding of incident management, problem management, root cause analysis, and operational support processes.
- Familiarity with source control, CI/CD pipelines, Infrastructure as Code, and DevOps practices.
- Strong analytical, problem-solving, communication, and stakeholder engagement skills.
- Preferred experience includes Git, Jenkins, Ansible, Terraform, Docker, Kubernetes, OpenShift, Kafka, Redis, MongoDB, Elasticsearch, AWS, Azure, and GCP.
Tech Stack
AnsibleApache KafkaAWSAzureC#C++DatadogDockerElasticsearchGitGoGoogle Cloud PlatformGrafanaJavaJenkinsKubernetesLinuxMongoDBOpenShiftPrometheusPythonRedisTerraform
Categories
DevOpsSite Reliability
About Jefferies
Jefferies is a global investment banking and capital markets firm serving corporations, financial institutions, governments, and individuals with M&A advisory, underwriting, sales and trading, research, and wealth management. Founded in 1962 and headquartered in New York, it operates from 40+ offices worldwide. The firm earns fees and trading revenues and is owned by publicly traded Jefferies Financial Group (NYSE: JEF).
