4 months ago
Remote, United KingdomStaff+

Responsibilities

  • Act as the technical lead for DevOps, platform, and release engineering by setting direction, standards, and best practices.
  • Architect and govern infrastructure provisioning, configuration management, CI/CD, release processes, and operations.
  • Design and support Windows-based high-availability solutions, including Windows clustering, failover patterns, maintenance, upgrades, and troubleshooting.
  • Lead Linux automation and platform standardization, including configuration, patching, hardening, and performance tuning.
  • Own Infrastructure as Code strategy with Terraform and automation strategy with Ansible.
  • Build and standardize deployments using Octopus Deploy, GitHub, and Ansible, including release promotion and rollback.
  • Design and mature CI/CD pipelines with artifact versioning, approvals, promotion strategies, and policy-as-code where applicable.
  • Establish observability standards using VictoriaMetrics and Prometheus for metrics, alerting, dashboards, and SLO/SLA monitoring.
  • Provide production leadership through incident response, root-cause analysis, postmortems, reliability improvements, and capacity planning.
  • Mentor engineers, review designs and code, lead cross-team initiatives, and maintain architecture documentation, runbooks, and platform roadmaps.

Requirements

  • Bachelor’s degree in Computer Science or a related field.
  • 7+ years or equivalent experience in DevOps, SRE, or infrastructure engineering, including leadership in complex environments.
  • Expert-level experience designing and operating Windows Server high-availability and clustering solutions.
  • Strong Linux administration and automation experience, including systemd, networking, storage, and performance.
  • Advanced Terraform and Ansible skills covering architecture, reusable components, and secure operations.
  • Strong deployment and release engineering experience with Octopus Deploy and GitHub.
  • Monitoring and observability expertise with VictoriaMetrics and/or Prometheus.
  • Production experience running Redis, RabbitMQ, and Nginx in high-availability environments, including tuning and troubleshooting.
  • Strong understanding of networking and security fundamentals, including TLS, DNS, load balancing, firewalling, and least privilege.
  • Proven ability to lead cross-team initiatives, make architectural decisions, and communicate clearly.
  • Experience with Kubernetes and container ecosystems, including Docker and Helm.
  • Preferred experience with GitLab CI, Jenkins, ELK/EFK, Loki, disaster recovery and business continuity, PowerShell, Vault, SOPS, VMware, Hyper-V, MS SQL Server, SIP, FreeSwitch, OpenSIP, session border controllers, F5 LTM, F5 GTM, AWS, or Azure.

Benefits

  • Primarily remote work with occasional visits to the Bristol office or London.
  • Opportunity to work on cloud communications and collaboration technology in a fast-paced, team-oriented environment.
  • Company culture emphasizes teamwork, transparency, accountability, and internal promotion.

Tech Stack

AnsibleAWSAzureDockerGitLab CI/CDHelmJenkinsKubernetesLinuxMicrosoft SQL ServerOctopus DeployPowerShellPrometheusRabbitMQRedisTerraformVault

Categories

Intermedia Intelligent Communications

About Intermedia Intelligent Communications

1,001-5,000 employees
Contact me