Base Salary
$113k - $172k/yr
Responsibilities
- Support and improve foundational networking, compute platforms, Kubernetes clusters, and ingress and traffic management systems.
- Harden existing systems and help roll out new infrastructure capabilities to improve platform reliability and scalability.
- Monitor system health using metrics, logs, and alerts.
- Participate in 24/7 on-call rotations to detect, respond to, and resolve incidents.
- Participate in standups, planning, retrospectives, and technical discussions while communicating progress and risks early.
- Stay current on technical trends and suggest innovative tools and approaches.
Requirements
- At least 3 years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles.
- Hands-on experience operating Linux-based systems in production.
- Working knowledge of networking fundamentals including load balancing, DNS, TLS, and ingress traffic flow.
- Experience with container orchestration such as EKS or Kubernetes.
- Experience with cloud-native infrastructure such as AWS, GCP, or Azure, including networking and compute concepts.
- Proficiency in at least one programming language such as Python, Ruby, or Go.
- Experience with Infrastructure as Code such as Terraform or CloudFormation.
- Preferred: experience with AWS cloud networking concepts including VPCs, subnets, routing, security groups, and load balancers.
- Preferred: experience operating or contributing to production Kubernetes platforms, including cluster upgrades, networking, or ingress configuration.
- Preferred: experience with monitoring, observability, and logging platforms such as DataDog, New Relic, SumoLogic, Splunk, Prometheus, or Grafana.
- Preferred: familiarity with service meshes, ingress controllers, or API gateways such as Envoy, Istio, or NGINX.
Benefits
- Hybrid work model based in established office locations, including Atlanta, with flexible work arrangements.
- Competitive salary and comprehensive benefits package.
- Company equity and Employee Stock Purchase Program, subject to eligibility.
- Retirement or pension plan, generous paid vacation, paid holidays, and sick leave.
- Dutonian Wellness Days and HibernationDuty companywide paid days off.
- Paid parental leave, including up to 22 weeks for a pregnant parent and 12 weeks for a non-pregnant parent, subject to local standards.
- 20 hours of paid volunteer time off per year.
- Companywide hack weeks and mental wellness programs.
Tech Stack
Categories
About PagerDuty
In an always-on world, teams trust PagerDuty to help them deliver an optimal digital experience to their customers, every time. PagerDuty is the central nervous system for a company’s digital operations. We identify issues and opportunities in real-time and bring together the right people to respond to problems faster and prevent them in the future. From digital disruptors to nearly half of the Fortune 500 and the Forbes AI 50s, over 34,000 paid and free customers rely on PagerDuty to help them continually improve their digital operations—so their teams can spend less time reacting to incidents and more time building for the future. Dutonians believe that we are a part of a bigger movement of businesses being built to benefit everyone—the customer and the employee, as well as our community. We are go-getters fueled by the fire to reinvent how people and companies work together. We take the lead and get creative to be first in the hearts of our customers. Whether it’s keeping the world on or changing it entirely, Dutonians are fueled by the fire to reinvent how people and companies work together to deliver in real-time, across the globe. Join us to lead uncharted efforts and reinvent how companies run.