Jefferies

Associate - AI Tooling Ops - Platform Reliability Engineer

Jefferies
Apply
4 days ago
Pune, IndiaMid Level

Responsibilities

  • Design, build, and maintain AI tooling infrastructure running on AWS Kubernetes.
  • Monitor platform health and improve system reliability, performance, scalability, and availability.
  • Perform incident triage, troubleshooting, stakeholder communication, and post-incident reviews.
  • Develop automation to reduce operational toil and manual support activities.
  • Build deployment, monitoring, alerting, dashboards, and observability capabilities.
  • Use Grafana, Datadog, Prometheus, and OpenTelemetry to monitor service health and analyze platform behavior.
  • Analyze logs, metrics, and distributed traces to identify bottlenecks and reliability issues.
  • Support Kafka-based enterprise messaging and event-driven architectures.
  • Troubleshoot applications, middleware, databases, messaging platforms, infrastructure, and cloud environments.
  • Participate in production support, problem management, release management, and change management activities.
  • Collaborate with global teams across APAC, EMEA, and the Americas on technology initiatives.

Requirements

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline.
  • At least 3 years of experience in Site Reliability Engineering, Platform Reliability Engineering, DevOps, Production Support, or Application Support.
  • Strong programming and scripting experience in Python, Go, C#, Java, or C++.
  • Understanding of software engineering principles, data structures, algorithms, and system design.
  • Experience supporting and troubleshooting distributed applications in production environments.
  • Working knowledge of Linux/Unix and Windows Server environments.
  • Experience with relational and NoSQL databases, including performance analysis and troubleshooting.
  • Hands-on experience with Grafana, Datadog, Prometheus, OpenTelemetry, Loki, and Jaeger.
  • Experience creating operational dashboards, alerts, runbooks, and monitoring solutions.
  • Understanding of event-driven architectures and enterprise messaging platforms such as Kafka.
  • Ability to troubleshoot message flows, APIs, middleware components, and distributed systems.
  • Understanding of incident management, problem management, root cause analysis, and operational support processes.
  • Familiarity with source control, CI/CD pipelines, Infrastructure as Code, and DevOps practices.
  • Strong analytical, problem-solving, communication, and stakeholder engagement skills.
  • Preferred experience includes Git, Jenkins, Ansible, Terraform, Docker, Kubernetes, OpenShift, Kafka, Redis, MongoDB, Elasticsearch, AWS, Azure, and GCP.

Tech Stack

AnsibleApache KafkaAWSAzureC#C++DatadogDockerElasticsearchGitGoGoogle Cloud PlatformGrafanaJavaJenkinsKubernetesLinuxMongoDBOpenShiftPrometheusPythonRedisTerraform

Categories

DevOpsSite Reliability
Jefferies

About Jefferies

5,001-10,000 employees

Jefferies is a global investment banking and capital markets firm serving corporations, financial institutions, governments, and individuals with M&A advisory, underwriting, sales and trading, research, and wealth management. Founded in 1962 and headquartered in New York, it operates from 40+ offices worldwide. The firm earns fees and trading revenues and is owned by publicly traded Jefferies Financial Group (NYSE: JEF).

Contact me