20 hours ago
Remote, United StatesSenior
Base Salary
$130k - $160k/yr
Responsibilities
- Own availability, performance, capacity, and reliability for production SaaS services in Azure, AWS, and FedRAMP High environments.
- Define and manage service-level indicators, service-level objectives, and error budgets with engineering and product teams.
- Build and tune Datadog and Azure Monitor dashboards, monitors, synthetic checks, anomaly detection, and alert routing.
- Automate remediation workflows and replace manual runbook procedures with code.
- Participate in on-call rotations, lead high-severity incident response, coordinate resolution, and complete post-incident reviews and customer-facing root cause analyses.
- Troubleshoot support escalations using logs, traces, network captures, and evidence-based issue routing.
- Build and maintain infrastructure as code with Terraform and Azure DevOps pipelines.
- Administer web application firewalls, including rule tuning, rate limiting, and false-positive triage.
- Manage observability costs involving metrics, APM, logs, indexes, retention, sampling, and archived storage.
- Operate within FedRAMP High change-control and continuous-monitoring processes and maintain runbooks and standard operating procedures.
- Improve on-call rotations, escalation paths, alert quality, runbook coverage, and regional handoffs.
- Partner with Support, Security, Product, and Development to ensure services launch with monitoring, runbooks, and SLOs.
Requirements
- 8+ years of experience in Site Reliability Engineering, DevOps, cloud operations, or production engineering for a SaaS product.
- Hands-on Azure experience with AKS, App Service, Azure SQL, Redis, Service Bus, Front Door, Storage, cloud networking, and cloud security fundamentals.
- Production experience with Datadog metrics, logs, APM, dashboards, monitor design, and query writing.
- Demonstrated ownership of SLIs, SLOs, and error budgets.
- Incident response experience, including leading incident bridges and writing postmortems.
- Production Kubernetes experience covering ingress, deployments, resource limits, and workload troubleshooting.
- Infrastructure as code experience with Terraform and CI/CD pipeline creation and troubleshooting, preferably with Azure DevOps.
- Scripting experience with PowerShell and Python, plus fluency with YAML and JSON.
- Strong networking and web fundamentals, including DNS, TLS, certificate chains, load balancing, reverse proxies, firewalls, and packet-level troubleshooting.
- Knowledge of cloud redundancy, backup, and disaster recovery strategies.
- Clear written communication for customer- and executive-facing root cause analyses.
- Willingness to participate in weekend and emergency on-call rotations.
- Preferred experience includes regulated environments such as FedRAMP, NIST 800-53, SOC 2, ISO 27001, or PCI; AWS with CloudFormation and SES; Jenkins; SaltStack; Consul; ELK; CloudWatch Logs Insights QL; Imperva; Azure WAF; Cloudflare; Microsoft Entra ID; SAML; OIDC; Jira Service Management; PagerDuty; multi-region architectures; disaster recovery testing; game days; chaos exercises; and mentoring engineers.
Benefits
- Competitive salary and meaningful bonus program.
- Healthcare insurance, pension or retirement matching, comprehensive life insurance, employee assistance program, time off plans, and paid company holidays.
- Up to 10% travel is required.
- The role includes participation in an on-call rotation covering weekends and emergencies.
Tech Stack
Categories
Site Reliability
About Delinea
Delinea builds identity security and privileged access management software for enterprises, covering human and machine identities across cloud and on‑premises environments. It sells a cloud‑native platform and related tools (such as PAM and privileged identity management) on a subscription basis to IT and security teams. Founded in 2021 in San Francisco through the merger of Thycotic and Centrify, it operates as a privately held company backed by TPG.
