NOV Inc.

Site Reliability Engineer

NOV Inc.
Apply
3 months ago
Houston, TX, USASenior

Responsibilities

  • Maintain and monitor production systems for availability, latency, and performance.
  • Lead incident response, including communication, resolution, postmortem documentation, and recurring-issue remediation.
  • Design and implement health checks, alerting systems, automated remediation workflows, and observability stacks.
  • Analyze telemetry and logs to identify trends, anomalies, and improvement opportunities.
  • Tune distributed systems, including AKKA.NET actors, and optimize PostgreSQL performance, queries, and maintenance strategies.
  • Collaborate with developers to improve architecture, throughput, latency, and stability.
  • Design and maintain CI/CD pipelines and automate deployment, testing, and rollback processes.
  • Standardize infrastructure-as-code practices across environments.

Requirements

  • 5+ years of experience in SRE, DevOps, or infrastructure engineering roles.
  • Expertise with Kubernetes and container orchestration at scale.
  • Strong experience with AKKA.NET or similar actor-based frameworks.
  • Proficiency in Bash, PowerShell, and Python scripting and automation.
  • Experience with Phobos, Datadog, Prometheus, Grafana, OpenTelemetry, or ELK.
  • Hands-on experience with AWS, Azure, or GCP.
  • Strong PostgreSQL knowledge, including performance tuning, query optimization, and maintenance.
  • Proven ability to lead incident management and postmortem processes.
  • Preferred experience with GitHub Actions, Azure Pipelines, GitLab CI, Docker, Terraform, GitHub, GitLab, Azure DevOps, C#, Python, Bash, and PowerShell.

Tech Stack

AWSAzureBashC#DatadogDockerGitHub ActionsGitLab CI/CDGoogle Cloud PlatformGrafanaKubernetesPostgreSQLPowerShellPrometheusPythonTerraform

Categories

Site Reliability
NOV Inc.

About NOV Inc.

10,000+ employees
Contact me