2 months ago
Remote, United StatesSenior
Responsibilities
- Build and maintain core platform infrastructure supporting engineering teams across the organization.
- Design and deliver AWS cloud infrastructure, containerized platforms, and cloud-native architectures using ECS and Kubernetes.
- Develop and operate CI/CD pipelines, infrastructure-as-code modules, reusable patterns, templates, and automation solutions.
- Create standardized cloud infrastructure patterns that support security, compliance, scalability, and rapid deployment.
- Improve software delivery practices, feedback cycles, observability, and operational excellence.
- Collaborate with product and engineering teams to refine requirements and deliver developer-productivity solutions.
- Troubleshoot complex distributed-system issues across networking, compute, storage, and application layers.
- Monitor and optimize cloud spend through cost governance and right-sizing recommendations.
- Maintain SLOs, SLIs, and SLAs; improve uptime and reliability engineering practices.
- Participate in incident management, on-call rotations, post-mortems, design reviews, and technical leadership.
- Mentor junior and mid-level engineers and promote knowledge sharing across the platform team.
- Work with emerging technologies including AI, MCP, and RAG.
Requirements
- At least 6 years of hands-on AWS experience focused on infrastructure as code.
- At least 4 years of experience managing and maintaining Linux systems.
- At least 3 years of Python experience for automation and scripting.
- At least 3 years of hands-on Terraform experience, including module development, state management, and multi-environment workflows.
- Practical experience with Git and automated workflows.
- Familiarity with AWS security practices, DNS, secure VPCs, and database security for PostgreSQL or MySQL.
- Experience troubleshooting distributed systems using packet captures, log analysis, performance profiling, and production-failure diagnosis.
- Knowledge of networking fundamentals including VPCs, load balancers, DNS, CDNs, service mesh, and network troubleshooting.
- Experience with monitoring and observability tools such as Datadog, Prometheus, Grafana, CloudWatch, and ELK stacks.
- Incident-management experience, including PagerDuty, on-call rotations, and post-mortem practices.
- Strong written and verbal communication skills and the ability to explain technical concepts to technical and non-technical audiences.
- Preferred qualifications include AWS certifications, experience with scalable application environments and event-driven architectures, Docker and Kubernetes security or orchestration, AI-related services, open-source or internal platform-tool contributions, AWS Organizations, Control Tower, Landing Zones, chaos engineering, game days, and Backstage.
Benefits
- External candidates may be required to attend an in-person interview at a VERSANT Media location before a hiring decision.
- The posting states that compensation may include health insurance, retirement plans, paid time off, and additional forms of compensation and benefits.
