
Senior Reliability Developer
Arctic Wolf1 month ago
Eden Prairie, MN, USASenior
Base Salary
$145k - $180k/yr
Responsibilities
- Manage changes and updates to public and private cloud infrastructure across AWS and OpenStack.
- Own monitoring for business services and customers, including dashboards, alerts, service health monitoring, and migration of legacy monitoring.
- Automate infrastructure provisioning, certificate lifecycle management, patching, and operational workflows using scripting languages and infrastructure-as-code tools.
- Write and manage Terraform modules and configuration trees for multi-region deployments.
- Administer Kubernetes clusters, containers, AWS services, EBS volumes, logging configurations, upgrades, scaling, and microservices environments.
- Maintain CI/CD pipelines and collaborate with development teams on features, fixes, troubleshooting, and production deployments.
- Monitor, renew, and redeploy TLS certificates and manage secrets using certificate and security tools.
- Troubleshoot distributed systems, API failures, service mesh architectures, container networking, inter-service dependencies, and service unavailability.
- Create and maintain MOPs, runbooks, service architecture diagrams, monitoring configurations, and other technical documentation.
- Plan and execute service decommissioning in laboratory and production environments.
- Participate in an on-call rotation, respond to escalations, and conduct root cause analysis during incident reviews.
Requirements
- Bachelor’s degree or foreign degree equivalent in Computer Information Systems or a related field, plus five years of progressive post-baccalaureate experience in a technology-related or related role.
- Experience designing, implementing, and managing multi-region cloud infrastructure with Terraform, Puppet, Chef, AWS, and OpenStack.
- Experience deploying and troubleshooting production microservices using Docker, Kubernetes, ECS, and EKS.
- Experience using Python, Bash, JavaScript, Groovy, and PowerShell for infrastructure automation, certificate management, and patching.
- Experience creating dashboards, configuring alerts, and implementing monitoring with Prometheus, Grafana, Zabbix, AlertManager, and PagerDuty.
- Experience architecting and managing highly available systems using AWS services including EC2, ECS, EKS, ELB, S3, RDS, IAM, Lambda, CloudFormation, VPC, Route53, and CloudWatch.
- Experience with Jenkins, GitLab, GitHub, Bitbucket, and Gaia for CI/CD automation and continuous integration workflows.
- Experience operating Kubernetes clusters across multiple AWS regions, including upgrades, logging, node scaling, EBS volume management, and pod orchestration.
- Experience with Keeper and TLS certificate monitoring, renewal coordination, and automated deployment.
- Experience troubleshooting distributed systems, microservices communications, API failures, service meshes, container networking, and deployment failures.
- Experience creating technical documentation with Jira, Confluence, Backstage, and LucidChart.
- Background checks are required, and the position may require authorization to access technology controlled under US export-control laws.
Benefits
- Telecommuting is permissible from any location within the United States.
- Equity is offered to all employees.
- Flexible time off and paid volunteer days are provided.
- RRSP and 401k matching are offered.
- Training and career development programs are available.
- Comprehensive private benefits include medical, mental health, dental, disability, life, AD&D, and value-added services.
- Employee Assistance Program services and fertility support are provided.
- Paid parental leave is offered.
- Background checks are required for the position.
Tech Stack
AWSBashChefDockerGrafanaGroovyJavaScriptJenkinsKubernetesOpenStackPostgreSQLPowerShellPrometheusPuppetPythonTerraform
Categories
DevOpsSite Reliability