Stord

Senior Site Reliability Engineer

Stord
Apply
2 months ago
Remote, United StatesSenior

Responsibilities

  • Own architecture and implementation of scalable, reliable GCP infrastructure across GKE, Cloud Run, AlloyDB, networking, and related services.
  • Develop and maintain Terraform modules, organization policies, reusable infrastructure patterns, and infrastructure provisioning automation.
  • Manage Kubernetes workloads, including performance tuning, capacity planning, resource optimization, and cost reduction.
  • Build monitoring, alerting, and observability capabilities in Datadog, including APM, logs, and RUM.
  • Develop disaster recovery and business-continuity strategies and validate their effectiveness.
  • Design and maintain GitHub Actions CI/CD pipelines, including runner strategy and deployment safety.
  • Create custom tooling and automate recurring operational workflows to reduce toil.
  • Partner with development and data teams to improve deployment practices and application reliability.
  • Provide escalation support for production incidents, participate in on-call, lead post-incident reviews, and implement durable fixes.
  • Participate in technical design reviews and improve SRE and infrastructure best practices across the team.

Requirements

  • 5+ years of experience in SRE, platform, or infrastructure engineering.
  • Strong hands-on experience with GCP core services, including GKE, Cloud Run, AlloyDB, networking, and IAM.
  • Fluency with Docker and Kubernetes, including debugging, tuning, and scaling production workloads.
  • Deep Terraform experience, including reusable modules and infrastructure state management.
  • Productive programming experience in TypeScript, Python, Go, or a similar language for tooling and automation.
  • Experience building actionable monitoring and alerting with Datadog or equivalent tools such as Prometheus and Grafana.
  • Understanding of distributed-systems fundamentals, including failure modes and consistency at scale.
  • Experience with Git, collaborative development workflows, incident management, and post-mortems.
  • Strong ownership, communication, collaboration, production mindset, and learning agility.
  • Ability to use AI coding tools productively while maintaining quality and technical judgment.
  • Preferred experience with PostgreSQL internals, database migrations and scaling, Redis, ClickHouse, analytical stores, Kafka, Redpanda, Pub/Sub, schema registries, cost engineering, GCP certifications, Cloudflare Workers, other Cloudflare services, or multi-cloud and hybrid architectures.

Benefits

  • High-autonomy work in a small, fast-moving SRE team with broad infrastructure ownership.
  • Opportunity to shape reliability, automation, cost-management, and infrastructure practices during significant company growth.
  • The role includes on-call participation for critical systems and hands-on production incident response.

Tech Stack

Apache KafkaClickHouseCloudflareDatadogDockerGitGitHub ActionsGoGoogle Cloud PlatformGrafanaKubernetesPostgreSQLPrometheusPythonRedisTerraformTypeScript

Categories

DevOpsSite Reliability
Stord

About Stord

1,001-5,000 employees
Contact me