2 hours ago
Hyderābād, IndiaStaff+
Responsibilities
- Support the operation and enhancement of mission-critical cloud-native environments.
- Deploy and manage application instrumentation for granular service-health insights.
- Help engineering teams implement and maintain application and infrastructure metrics, logs, and traces.
- Unify observability tooling and centralize telemetry through platforms such as Application Insights.
- Configure OpenTelemetry-based collection within Azure Monitor Application Insights.
- Apply semantic conventions to AI-agent observability data and support structured logging, distributed tracing, and core metrics.
- Support incident response and root-cause analysis using logs, metrics, traces, and observability tools.
- Build automation to reduce toil, improve developer experience, engineering velocity, and system reliability.
- Define and manage SLOs and error budgets with engineering teams.
- Participate in rotational evening shifts and provide weekend or on-call support as needed.
- Collaborate with Agile teams and participate in design discussions with clients, vendors, and stakeholders.
- Share knowledge across product areas and apply ITIL practices for incident, change, and problem management.
Requirements
- Bachelor’s degree in Computer Science or a related field; a master’s degree is a plus.
- At least 5 years of experience in Site Reliability, Observability, DevOps, or Cloud Engineering roles.
- Expertise with Microsoft Azure Cloud and observability frameworks such as OpenTelemetry and distributed tracing systems.
- Experience with infrastructure as code using Bicep, ARM, and Terraform.
- Strong understanding of instrumenting, tracing, and correlating AI/LLM workflows with infrastructure telemetry.
- Experience with Azure Monitor, Application Insights, DataDog, and Log Analytics.
- Knowledge of AI/ML-based anomaly detection and log aggregation and analysis tools such as Microsoft Azure Anomaly Detector.
- Experience with agentic or LLM-based systems, including LangChain, Celery, OpenAI APIs, and orchestration frameworks.
- Experience with application reliability platforms such as Checkly and synthetic monitoring using Playwright.
- Understanding of networking, Kubernetes, Docker, APIs, scripting languages, and databases including SQL, Cosmos DB, and PostgreSQL.
- Familiarity with SimCorp Dimension and Salesforce is a plus.
- Proficiency with IT service management frameworks such as ITIL.
- Experience managing onboarding projects and live production operations.
- Ability to work collaboratively in cross-functional teams and interest in continuous learning.
Benefits
- Global hybrid work policy with two required office days per week and remote work on other days.
- Inclusive and diverse company culture.
- Work-life balance focused on equilibrium between professional and personal responsibilities.
- Empowerment through participation in shaping work processes.
- Professional development and individualized career-growth opportunities.
Tech Stack
Categories
DevOpsSite Reliability
