
Cloud Infrastructure Engineer (Open LMS) UK, Remote
Lietuvos Geležinkeliai1 month ago
Remote, United KingdomSenior
Responsibilities
- Design, build, and maintain AWS infrastructure using Terraform across services including EC2, RDS, S3, SQS, Lambda, ALB, ElastiCache, Route 53, and VPC networking.
- Write and maintain Puppet modules for fleets of EC2 instances across multiple auto-scaling groups.
- Maintain and extend Python-based automation, tooling, and production daemons supporting platform operations.
- Operate distributed service discovery and configuration management with etcd.
- Manage and tune caching using Varnish, Redis/Valkey, and PHP OPcache.
- Run and scale the Prometheus, Grafana, Loki, Fluentd, and PagerDuty observability stack.
- Evaluate and implement distributed storage solutions as the platform evolves.
- Improve deployment workflows and release processes for non-containerized environments.
- Collaborate on API contracts, integration patterns, and operational tooling.
- Participate in on-call rotations, incident response, root cause analysis, and platform reliability improvements.
Requirements
- Strong production experience with AWS, particularly EC2, RDS, S3, SQS, Lambda, ALB, ElastiCache, Route 53, IAM, and VPC networking.
- Proficiency authoring and maintaining Terraform modules for production infrastructure.
- Proficiency authoring and maintaining Puppet modules or equivalent agent-based configuration management.
- Solid Python skills for production automation and daemons.
- Deep Linux systems knowledge, particularly Ubuntu, with familiarity with Apache, Nginx, PHP-FPM, Varnish, systemd, filesystem mounts, and networking fundamentals.
- Understanding of distributed systems concepts including consensus, leader election, distributed locking, and eventual consistency.
- Production experience building and maintaining observability pipelines using Prometheus, Grafana, Loki, or equivalent tools.
- Comfort working in a GitLab-based CI/CD workflow.
- Ability to document architectural decisions and explain technical tradeoffs to technical and non-technical stakeholders.
- Preferred experience with distributed storage systems such as Ceph, GlusterFS, JuiceFS, CubeFS, or AWS EFS.
- Preferred familiarity with etcd, Consul, or ZooKeeper, including watch APIs, TTL-based locking, and cluster operations.
- Preferred experience with Varnish and VCL, multi-tenant SaaS platform design, PHP, Moodle LMS, secrets management, credential rotation, and zero-downtime VM-based deployments.