2 months ago
Remote, United StatesSenior
Responsibilities
- Own architecture and implementation of scalable, reliable GCP infrastructure across GKE, Cloud Run, AlloyDB, networking, and related services.
- Develop and maintain Terraform modules, organization policies, reusable infrastructure patterns, and infrastructure provisioning automation.
- Manage Kubernetes workloads, including performance tuning, capacity planning, resource optimization, and cost reduction.
- Build monitoring, alerting, and observability capabilities in Datadog, including APM, logs, and RUM.
- Develop disaster recovery and business-continuity strategies and validate their effectiveness.
- Design and maintain GitHub Actions CI/CD pipelines, including runner strategy and deployment safety.
- Create custom tooling and automate recurring operational workflows to reduce toil.
- Partner with development and data teams to improve deployment practices and application reliability.
- Provide escalation support for production incidents, participate in on-call, lead post-incident reviews, and implement durable fixes.
- Participate in technical design reviews and improve SRE and infrastructure best practices across the team.
Requirements
- 5+ years of experience in SRE, platform, or infrastructure engineering.
- Strong hands-on experience with GCP core services, including GKE, Cloud Run, AlloyDB, networking, and IAM.
- Fluency with Docker and Kubernetes, including debugging, tuning, and scaling production workloads.
- Deep Terraform experience, including reusable modules and infrastructure state management.
- Productive programming experience in TypeScript, Python, Go, or a similar language for tooling and automation.
- Experience building actionable monitoring and alerting with Datadog or equivalent tools such as Prometheus and Grafana.
- Understanding of distributed-systems fundamentals, including failure modes and consistency at scale.
- Experience with Git, collaborative development workflows, incident management, and post-mortems.
- Strong ownership, communication, collaboration, production mindset, and learning agility.
- Ability to use AI coding tools productively while maintaining quality and technical judgment.
- Preferred experience with PostgreSQL internals, database migrations and scaling, Redis, ClickHouse, analytical stores, Kafka, Redpanda, Pub/Sub, schema registries, cost engineering, GCP certifications, Cloudflare Workers, other Cloudflare services, or multi-cloud and hybrid architectures.
Benefits
- High-autonomy work in a small, fast-moving SRE team with broad infrastructure ownership.
- Opportunity to shape reliability, automation, cost-management, and infrastructure practices during significant company growth.
- The role includes on-call participation for critical systems and hands-on production incident response.
Tech Stack
Apache KafkaClickHouseCloudflareDatadogDockerGitGitHub ActionsGoGoogle Cloud PlatformGrafanaKubernetesPostgreSQLPrometheusPythonRedisTerraformTypeScript
Categories
DevOpsSite Reliability
