
Principal Site Reliability Engineer
Early Warning Services2 hours ago
Scottsdale, AZ, USAStaff+
Base Salary
$194k - $237k/yr
Responsibilities
- Design and implement software and tools that improve application and infrastructure availability, scalability, performance, and latency.
- Build automation and tooling for application management, deployments, configuration changes, and disaster recovery.
- Design and promote observability and monitoring systems to detect problems proactively and identify root causes.
- Continuously evaluate application capacity, provide statistics to product and business teams, and recommend scaling strategies.
- Identify performance bottlenecks and troubleshoot technical issues with cross-functional teams.
- Lead high-priority incident response and improve the root-cause analysis process.
- Serve as a technical liaison, creating documentation and runbooks for Level 1 and Level 2 teams.
- Participate in a 24/7 on-call rotation and support customer-facing production applications.
- Partner with application development teams on microservice design patterns, software reliability practices, and software development lifecycle requirements.
- Lead cross-functional technical solutions, mentor team members, and establish repeatable, reusable engineering practices.
Requirements
- Bachelor’s degree in business, computer science, or a related field.
- 12+ years of related experience managing large, complex projects in a technical or software development environment, inclusive of post-graduate degree.
- Proven ability to lead teams through high-priority incidents and improve root-cause analysis processes.
- Excellent troubleshooting skills and experience resolving technical issues in complex environments.
- Hands-on development experience with one or more of Python, Go, or Java.
- Experience with Docker, microservices architecture, messaging frameworks, database technologies, and caching layers.
- Strong understanding of Linux administration and networking fundamentals.
- Experience implementing CI/CD pipelines using tools such as Git, Chef, Maven, and Jenkins.
- Experience leading cross-functional teams and designing and building complex end-to-end systems.
- Preferred: programming experience with Java, Ruby, Python, JavaScript, or Go.
- Preferred: experience supporting applications in a 24/7 customer-facing production environment.
- Preferred: working knowledge of AWS, Docker, Kubernetes, and Swarm.
- Successful completion of a background and drug screen is required.
Benefits
- Hybrid work model for positions in Scottsdale, San Francisco, Chicago, or New York.
- Medical, dental, and vision coverage, including PPO and HDHP options, with HSA and FSA contributions or savings options.
- 401(k) retirement plan with a 100% company Safe Harbor Match on the first 6% of employee deferral immediately upon eligibility.
- Flexible Time Off for exempt employees, generous PTO for non-exempt employees, 11 paid company holidays, and one paid volunteer day.
- 12 weeks of paid parental leave.
- Maven Family Planning support covering egg freezing, fertility, adoption, surrogacy, pregnancy, postpartum, early pediatrics, and return-to-work support.
- Normal office environment with primarily sedentary computer-based work and reasonable accommodation available when needed.
Tech Stack
Amazon DynamoDBApache KafkaAWSChefDockerGitGoJavaJavaScriptJenkinsKubernetesLinuxMavenPythonRedisRuby
Categories
Site Reliability