Goods & Services

Senior DevOps/Platform Engineer

Goods & Services
Apply
4 days ago
Remote, MexicoSenior

Responsibilities

  • Provision and maintain development, test, and production environments, access management, CI/CD pipelines, and application-delivery infrastructure.
  • Implement and maintain platform security and observability capabilities, including WAF, rate limiting, distributed tracing, monitoring, dashboards, alerting, and paging workflows.
  • Build and maintain feature-flag infrastructure for controlled releases, incremental migrations, parallel runs, and safe rollbacks.
  • Monitor event-driven and asynchronous infrastructure, including event failures, message queues, dead-letter queues, schema lifecycle issues, and backup or archival processes.
  • Develop service reliability and integration monitoring for internal and third-party services, including failure detection, dashboards, circuit-breaker monitoring, and automated notifications.
  • Create operational runbooks and incident-response procedures, troubleshoot production issues, and define recovery procedures.
  • Support staged production rollouts and traffic validation, including ramp-up criteria, soak periods, stabilization periods, validation metrics, and rollback strategies.
  • Partner with engineering and technical leadership to establish deployment standards, operational processes, and production-readiness criteria.

Requirements

  • Solid AWS CI/CD experience with tools such as GitHub Actions, AWS CodePipeline, or equivalent in multi-service cloud environments.
  • Hands-on experience with AWS observability and monitoring, including CloudWatch, AWS X-Ray, logging, alerting, dashboards, and paging workflows.
  • Experience securing and hardening AWS WAF and API Gateway, including managed security rules, rate limiting, access controls, and traffic protection.
  • Experience implementing and managing feature-flag platforms for controlled releases, incremental migrations, parallel-run scenarios, and safe rollbacks.
  • Experience with infrastructure and environment management, including provisioning, configuration, access management, and support for development, test, and production environments.
  • Experience monitoring event-driven and asynchronous systems, including EventBridge, message queues, dead-letter queues, event failures, and alerting mechanisms.
  • Experience implementing service reliability and integration monitoring, including circuit-breaker monitoring, failure detection, dashboards, and automated alerting for third-party services.
  • Solid incident-response and operational-readiness experience, including creating runbooks, troubleshooting production issues, and defining recovery procedures.
  • Experience supporting production deployments and controlled traffic rollouts, including ramp-up criteria, soak periods, validation metrics, and rollback strategies.
  • Ability to collaborate with engineering and technical leadership to establish deployment standards, operational processes, and production-readiness criteria.

Tech Stack

AWSGitHub Actions

Categories

Goods & Services

About Goods & Services

201-500 employees
Contact me