SimCorp

Lead Site Reliability Engineer - Observability

SimCorp
Apply
2 hours ago
Hyderābād, IndiaStaff+

Responsibilities

  • Support the operation and enhancement of mission-critical cloud-native environments.
  • Deploy and manage application instrumentation for granular service-health insights.
  • Help engineering teams implement and maintain application and infrastructure metrics, logs, and traces.
  • Unify observability tooling and centralize telemetry through platforms such as Application Insights.
  • Configure OpenTelemetry-based collection within Azure Monitor Application Insights.
  • Apply semantic conventions to AI-agent observability data and support structured logging, distributed tracing, and core metrics.
  • Support incident response and root-cause analysis using logs, metrics, traces, and observability tools.
  • Build automation to reduce toil, improve developer experience, engineering velocity, and system reliability.
  • Define and manage SLOs and error budgets with engineering teams.
  • Participate in rotational evening shifts and provide weekend or on-call support as needed.
  • Collaborate with Agile teams and participate in design discussions with clients, vendors, and stakeholders.
  • Share knowledge across product areas and apply ITIL practices for incident, change, and problem management.

Requirements

  • Bachelor’s degree in Computer Science or a related field; a master’s degree is a plus.
  • At least 5 years of experience in Site Reliability, Observability, DevOps, or Cloud Engineering roles.
  • Expertise with Microsoft Azure Cloud and observability frameworks such as OpenTelemetry and distributed tracing systems.
  • Experience with infrastructure as code using Bicep, ARM, and Terraform.
  • Strong understanding of instrumenting, tracing, and correlating AI/LLM workflows with infrastructure telemetry.
  • Experience with Azure Monitor, Application Insights, DataDog, and Log Analytics.
  • Knowledge of AI/ML-based anomaly detection and log aggregation and analysis tools such as Microsoft Azure Anomaly Detector.
  • Experience with agentic or LLM-based systems, including LangChain, Celery, OpenAI APIs, and orchestration frameworks.
  • Experience with application reliability platforms such as Checkly and synthetic monitoring using Playwright.
  • Understanding of networking, Kubernetes, Docker, APIs, scripting languages, and databases including SQL, Cosmos DB, and PostgreSQL.
  • Familiarity with SimCorp Dimension and Salesforce is a plus.
  • Proficiency with IT service management frameworks such as ITIL.
  • Experience managing onboarding projects and live production operations.
  • Ability to work collaboratively in cross-functional teams and interest in continuous learning.

Benefits

  • Global hybrid work policy with two required office days per week and remote work on other days.
  • Inclusive and diverse company culture.
  • Work-life balance focused on equilibrium between professional and personal responsibilities.
  • Empowerment through participation in shaping work processes.
  • Professional development and individualized career-growth opportunities.

Tech Stack

Categories

DevOpsSite Reliability
SimCorp

About SimCorp

1,001-5,000 employees
Contact me