
Sr Lead Site Reliability Engineer
Lumen Technologies3 days ago
Remote, United StatesStaff+
Base Salary
$132k - $194k/yr
Responsibilities
- Provide production support and improve incident management processes for customer-facing applications.
- Implement AI-assisted outage triage, analysis, remediation, and root-cause analysis.
- Monitor and optimize application, database, and infrastructure performance, including latency, throughput, and resource utilization.
- Deploy and maintain monitoring, alerting, dashboards, metrics, anomaly detection, and noise-reduction capabilities.
- Develop automation for deployment, scaling, and recovery using AI-assisted code generation and validation tools.
- Manage AWS resources with Terraform or similar infrastructure-as-code tools.
- Analyze system dependencies and implement improvements to availability, resilience, and fault tolerance.
- Champion SLIs, SLOs, error budgets, resilient architecture, and reliability improvements for the Lumen Connect platform.
- Collaborate with software engineering, DevOps, product, development, and testing teams on reliability goals.
- Document processes, runbooks, and best practices and provide mentorship on reliability and operational excellence.
Requirements
- 10+ years of professional experience in SRE, DevOps, or infrastructure engineering roles.
- Experience with Terraform or similar infrastructure-as-code tools for cloud resource management.
- Proficiency with scripting languages such as Python and Bash and with automation frameworks.
- Experience with CI/CD pipelines and tools such as GitHub Actions, Jenkins, or GitLab CI.
- Understanding of monitoring and logging tools such as CloudWatch, ELK, and Datadog.
- Familiarity with Docker and Kubernetes.
- Strong AI and problem-solving skills with a proactive mindset.
- Preferred experience with AWS services including EC2, CloudFront, EKS, RDS, and S3.
- AWS or related technology certifications are preferred.
- Preferred application development experience with Java microservices and Spring Boot.
- Experience with Agile/SCRUM methodologies and development practices is preferred.
Benefits
- Fully remote position within the United States.
- Comprehensive health, life, voluntary lifestyle, and other benefits supporting physical, mental, emotional, and financial wellbeing.
- Bonus structure information, including short-term incentives, long-term incentives, and/or sales compensation, is available during the selection process.
Tech Stack
Categories
Site Reliability