3 months ago
Doha, QatarSenior
Responsibilities
- Own infrastructure for large-scale, distributed production systems in multi-account AWS environments.
- Operate ECS and/or Kubernetes workloads, including capacity providers, task definitions, service autoscaling, and production orchestration.
- Manage production databases, including schema migrations, performance tuning, connection management, and failover.
- Design and maintain CI/CD pipelines across multi-team and multi-environment delivery workflows.
- Operate observability systems covering metrics, logging, tracing, and alerting, including cost optimization at scale.
- Develop modular infrastructure-as-code, manage state, detect drift, and execute multi-account rollouts with Terraform or Pulumi.
- Build internal platforms and automation using Python and/or Go.
- Diagnose production networking and connectivity issues involving DNS, load balancing, security groups, and VPNs.
- Build compliance controls into infrastructure under PCI-DSS, SOC 2, and ISO 27001 frameworks.
- Partner with engineering leadership to connect infrastructure decisions to availability, cost efficiency, and delivery velocity.
Requirements
- At least 6 years of experience in a DevOps, CloudOps, or SRE role owning infrastructure for large-scale distributed production systems.
- Deep AWS expertise across multi-account environments, including Organizations, IAM, VPC, ECS, RDS, Secrets Manager, and cost management tooling.
- Advanced hands-on production experience with ECS and/or Kubernetes.
- Production-grade operations experience with PostgreSQL and/or other relational or NoSQL database engines.
- Proven experience designing and maintaining GitHub Actions CI/CD pipelines.
- Fluency with Datadog, Prometheus/Grafana, OpenSearch, or comparable observability tooling.
- Advanced Terraform or Pulumi infrastructure-as-code experience covering modular design, state management, drift detection, and multi-account rollouts.
- Strong automation and tooling development skills in Python and/or Go, including internal platform development.
- Solid networking fundamentals and the ability to diagnose production connectivity issues without a runbook.
- Experience operating under PCI-DSS, SOC 2, and ISO 27001 compliance frameworks.
- Ability to work with engineering leadership and translate infrastructure decisions into business outcomes.
Benefits
- Global collaboration with a worldwide team.
- Learning budgets and access to courses and growth tools.
- Autonomy and ownership over tasks and career path.
- Flexible time off, generous leave, and wellness policies.
- A Scrum-oriented work environment.
- Great Place to Work®-certified workplace with an emphasis on inclusion and employee wellbeing.
