GE Vernova

SRE Observability SLO Engineer

GE Vernova
Apply
12 days ago
Mexico City, Mexico or Monterrey, MexicoMid Level / Senior

Responsibilities

  • Implement telemetry standards for metrics, logs, and distributed traces across GridOS SaaS services.
  • Implement Kubernetes metrics collection and define retention, cardinality, and telemetry cost controls.
  • Build and maintain observability runbooks, SLO tooling, error-budget alerts, dashboards, and compliance reports.
  • Define SLIs and SLOs with engineering and customer stakeholders and coordinate SLO reviews and reliability prioritization.
  • Design operational, executive, and customer-facing dashboards covering availability, latency, errors, and saturation.
  • Implement alert policies, routing, escalation procedures, and on-call schedules.
  • Design synthetic monitoring for critical user journeys, APIs, user interfaces, and integration endpoints.
  • Continuously expand observability coverage, improve alert signal-to-noise, reduce MTTD and MTTR, and validate infrastructure costs against reliability data.

Requirements

  • Bachelor's degree in Computer Science.
  • 2–3 years of experience in SRE, observability engineering, or infrastructure reliability roles.
  • Fluency in English.
  • Experience with at least one major observability platform, such as Datadog, Grafana with Prometheus, AWS CloudWatch, Dynatrace, or New Relic.
  • Understanding of distributed telemetry, including metrics, structured logging, and distributed tracing.
  • Experience with Kubernetes observability, kube-state-metrics, node exporters, Helm-deployed monitoring stacks, and namespace-level resource metrics.
  • Proficiency with at least one query or visualization language, including PromQL, Splunk SPL, Datadog Query Language, or CloudWatch Logs Insights query syntax.
  • Experience configuring monitoring alerts for system health visibility.
  • Scripting skills in Python and/or Bash.
  • Familiarity with AWS and Kubernetes infrastructure, deployment and configuration tools, observability platforms, alerting systems, and Linux administration.
  • Familiarity with OpenTelemetry, synthetic monitoring tools, and regulated-industry compliance requirements is preferred.
  • AWS certifications in CloudWatch/Observability or Solutions Architecture are preferred.

Benefits

  • Relocation assistance is provided.

Tech Stack

AnsibleAWSAzureBashChefDatadogGoGrafanaGroovyHelmKubernetesLinuxPrometheusPuppetPythonRancherSplunk

Categories

Site Reliability
GE Vernova

About GE Vernova

10,000+ employees

GE Vernova is a purpose-built energy technology company on a mission to electrify to thrive and decarbonize the world. It is made up of three businesses -- Power, Electrification, and Wind -- with focus on accelerating the path to more reliable, affordable, and sustainable energy, while helping our customers power economies and deliver the electricity that is vital to health, safety, security, and improved quality of life. The world needs more energy, smarter energy. With energy demand expected to grow by more than 50% in the next 20 years, we are continuously innovating to meet the moment…like we have for the past 130 years. The Energy of Change and relentless optimism are what drive us – it’s about never giving up and seeing what’s possible so that we deliver the energy technologies the world needs right now and for generations to come. GE Vernova’s attitude and edge is embedded in its name. We retain our treasured legacy, “GE,” as an enduring and hard-earned badge of quality and ingenuity. “Ver” / “verde” signal Earth’s verdant and lush ecosystems. “Nova,” from the Latin “novus,” nods to a new, innovative era of lower carbon energy that GE Vernova will help deliver. Together, we have the energy to change the world.

Contact me