3 months ago
Houston, TX, USASenior
Responsibilities
- Maintain and monitor production systems for availability, latency, and performance.
- Lead incident response, including communication, resolution, postmortem documentation, and recurring-issue remediation.
- Design and implement health checks, alerting systems, automated remediation workflows, and observability stacks.
- Analyze telemetry and logs to identify trends, anomalies, and improvement opportunities.
- Tune distributed systems, including AKKA.NET actors, and optimize PostgreSQL performance, queries, and maintenance strategies.
- Collaborate with developers to improve architecture, throughput, latency, and stability.
- Design and maintain CI/CD pipelines and automate deployment, testing, and rollback processes.
- Standardize infrastructure-as-code practices across environments.
Requirements
- 5+ years of experience in SRE, DevOps, or infrastructure engineering roles.
- Expertise with Kubernetes and container orchestration at scale.
- Strong experience with AKKA.NET or similar actor-based frameworks.
- Proficiency in Bash, PowerShell, and Python scripting and automation.
- Experience with Phobos, Datadog, Prometheus, Grafana, OpenTelemetry, or ELK.
- Hands-on experience with AWS, Azure, or GCP.
- Strong PostgreSQL knowledge, including performance tuning, query optimization, and maintenance.
- Proven ability to lead incident management and postmortem processes.
- Preferred experience with GitHub Actions, Azure Pipelines, GitLab CI, Docker, Terraform, GitHub, GitLab, Azure DevOps, C#, Python, Bash, and PowerShell.
Tech Stack
AWSAzureBashC#DatadogDockerGitHub ActionsGitLab CI/CDGoogle Cloud PlatformGrafanaKubernetesPostgreSQLPowerShellPrometheusPythonTerraform
Categories
Site Reliability
