Principal DevOps Engineer
Intermedia Intelligent Communications4 months ago
Remote, United KingdomStaff+
Responsibilities
- Act as the technical lead for DevOps, platform, and release engineering by setting direction, standards, and best practices.
- Architect and govern infrastructure provisioning, configuration management, CI/CD, release processes, and operations.
- Design and support Windows-based high-availability solutions, including Windows clustering, failover patterns, maintenance, upgrades, and troubleshooting.
- Lead Linux automation and platform standardization, including configuration, patching, hardening, and performance tuning.
- Own Infrastructure as Code strategy with Terraform and automation strategy with Ansible.
- Build and standardize deployments using Octopus Deploy, GitHub, and Ansible, including release promotion and rollback.
- Design and mature CI/CD pipelines with artifact versioning, approvals, promotion strategies, and policy-as-code where applicable.
- Establish observability standards using VictoriaMetrics and Prometheus for metrics, alerting, dashboards, and SLO/SLA monitoring.
- Provide production leadership through incident response, root-cause analysis, postmortems, reliability improvements, and capacity planning.
- Mentor engineers, review designs and code, lead cross-team initiatives, and maintain architecture documentation, runbooks, and platform roadmaps.
Requirements
- Bachelor’s degree in Computer Science or a related field.
- 7+ years or equivalent experience in DevOps, SRE, or infrastructure engineering, including leadership in complex environments.
- Expert-level experience designing and operating Windows Server high-availability and clustering solutions.
- Strong Linux administration and automation experience, including systemd, networking, storage, and performance.
- Advanced Terraform and Ansible skills covering architecture, reusable components, and secure operations.
- Strong deployment and release engineering experience with Octopus Deploy and GitHub.
- Monitoring and observability expertise with VictoriaMetrics and/or Prometheus.
- Production experience running Redis, RabbitMQ, and Nginx in high-availability environments, including tuning and troubleshooting.
- Strong understanding of networking and security fundamentals, including TLS, DNS, load balancing, firewalling, and least privilege.
- Proven ability to lead cross-team initiatives, make architectural decisions, and communicate clearly.
- Experience with Kubernetes and container ecosystems, including Docker and Helm.
- Preferred experience with GitLab CI, Jenkins, ELK/EFK, Loki, disaster recovery and business continuity, PowerShell, Vault, SOPS, VMware, Hyper-V, MS SQL Server, SIP, FreeSwitch, OpenSIP, session border controllers, F5 LTM, F5 GTM, AWS, or Azure.
Benefits
- Primarily remote work with occasional visits to the Bristol office or London.
- Opportunity to work on cloud communications and collaboration technology in a fast-paced, team-oriented environment.
- Company culture emphasizes teamwork, transparency, accountability, and internal promotion.