StoneX Group Inc.

Lead Engineer - Reliability Engineering

StoneX Group Inc.
Apply
2 hours ago
Bengaluru, IndiaStaff+

Responsibilities

  • Define and drive reliability engineering standards, practices, maturity models, service tiering, and adoption metrics across platforms and services.
  • Partner with engineering, platform, infrastructure, and product teams to improve reliability, resilience, operability, supportability, and service ownership.
  • Establish SLOs, SLIs, error budgets, alert-quality standards, toil-reduction practices, production-readiness reviews, and service ownership expectations.
  • Improve change safety, release confidence, capacity planning, resilience testing, disaster recovery, incident management, post-incident reviews, and on-call effectiveness.
  • Drive end-to-end observability adoption, including metrics, logs, traces, service maps, dashboards, instrumentation, actionable alerting, and service health reporting.
  • Define observability architecture and telemetry standards and embed reliability guardrails into CI/CD pipelines, internal developer platforms, templates, workflows, scorecards, and dashboards.
  • Identify and address reliability risks, architectural weaknesses, service fragility, and third-party or provider dependencies.
  • Define and track reliability metrics and operational KPIs that improve service outcomes over time.
  • Contribute hands-on to technical design, implementation, and operational improvement while guiding technical direction.
  • Mentor and coach engineers and help raise reliability maturity across the organization.

Requirements

  • 7+ years of experience in SRE, production engineering, platform engineering, infrastructure engineering, or a closely related role.
  • Track record of building, improving, or scaling reliability engineering or SRE practices in a large organization.
  • Several years of hands-on experience supporting production systems at scale, including incident response, problem management, availability improvement, and operational excellence.
  • Strong experience defining and implementing SLOs, SLIs, error budgets, service health models, and reliability-focused engineering practices.
  • Strong experience with observability platforms such as Datadog, including metrics, logs, tracing, alerting, dashboards, and service-level reporting.
  • Experience driving end-to-end observability adoption across application teams, including instrumentation, telemetry standards, dashboards, alerting, and service-level reporting.
  • Experience automating operational processes with Terraform, scripting languages, CI/CD pipelines, and cloud-native platforms.
  • Experience with Kubernetes, Linux, Git, and modern cloud or platform infrastructure.
  • Strong systems thinking and the ability to balance reliability, latency, engineering velocity, risk, and cost.
  • Strong communication, collaboration, influencing, mentoring, and cross-functional leadership skills.
  • Bachelor’s degree in computer science, engineering, or a related field, or equivalent practical experience.
  • Relevant certifications are a plus, along with a commitment to continual professional and technical development.

Benefits

  • Four days’ work from the office.
  • Occasional travel for team collaboration meetings and conferences.

Categories

Site Reliability
StoneX Group Inc.

About StoneX Group Inc.

5,001-10,000 employees
Contact me