1 hour ago
Remote, India or Bengaluru, IndiaMid Level / Senior

Responsibilities

  • Build, operate, maintain, and enhance enterprise observability platforms for applications and infrastructure.
  • Operate observability backends, ingestion pipelines, agents, collectors, and data lifecycle processes.
  • Build and maintain telemetry pipelines for logs, metrics, traces, and events.
  • Lead telemetry onboarding for applications, infrastructure, Kubernetes, cloud, and platform teams.
  • Provide internal product support by troubleshooting ingestion, query performance, data quality, and platform usage issues.
  • Improve platform availability, performance, capacity, scalability, and reliability.
  • Troubleshoot production issues involving telemetry ingestion, processing, storage, indexing, and query performance.
  • Optimize data stores, retention policies, indexing strategies, and storage utilization.
  • Build and maintain dashboards, alerts, integrations, and operational tooling.
  • Automate platform provisioning, configuration, and deployments using Infrastructure as Code and CI/CD.
  • Partner with Application, SRE, Infrastructure, and Security teams to improve telemetry quality, coverage, and incident response.

Requirements

  • At least 7 years of hands-on experience in Platform Engineering, DevOps, SRE, Infrastructure Engineering, or Observability Engineering.
  • Strong hands-on experience with at least one major observability stack such as Elastic, Grafana, Datadog, Splunk, or an equivalent.
  • Experience operating and troubleshooting production observability platforms at scale.
  • Strong understanding of logs, metrics, distributed tracing, telemetry pipelines, and observability concepts.
  • Experience onboarding telemetry and data sources and supporting internal customers with integration and troubleshooting.
  • Hands-on experience with Kubernetes, containers, Linux, networking, and cloud infrastructure.
  • Familiarity with OpenTelemetry or similar instrumentation and collection frameworks.
  • Advanced troubleshooting experience across applications, infrastructure, and distributed systems.
  • Hands-on experience with Infrastructure as Code tools such as Terraform.
  • Experience building CI/CD pipelines using GitHub Actions, Azure DevOps, Jenkins, or similar tools.
  • Hands-on automation and scripting experience using Python, Go, Bash, or similar languages.
  • Experience with performance tuning, capacity planning, data lifecycle management, and platform optimization.
  • Knowledge of telemetry collection, signal processing, cross-signal analysis, alerting, and production troubleshooting.
  • Ability to build, operate, troubleshoot, and improve production-grade platform solutions.
  • Experience implementing SLIs, SLOs, alerting standards, and reliability monitoring.
  • Preferred experience with Grafana, Prometheus, ClickHouse, pipeline or stream processing platforms, Fluent Bit, Kafka, high-volume telemetry environments, cost optimization, GitOps workflows, observability-as-code, onboarding documentation, runbooks, and self-service guidance.

Benefits

  • The company states that it is committed to an inclusive, respectful, and supportive work environment.
  • First American (India) is an Equal Opportunity Employer and states that it does not discriminate based on protected characteristics.

Tech Stack

Apache KafkaBashClickHouseDatadogGitHub ActionsGoGrafanaJenkinsKubernetesLinuxPrometheusPythonSplunkTerraform

Categories

First American Financial Corporation

About First American Financial Corporation

10,000+ employees
Contact me