
Staff Site Reliability Engineer
SoundHound, Inc.1 month ago
Remote, Canada or Toronto, CanadaStaff+
Responsibilities
- Design, build, and maintain highly available and scalable infrastructure on Google Cloud Platform.
- Architect and automate CI/CD pipelines for rapid and reliable deployments.
- Implement monitoring, alerting, and observability strategies to identify and resolve system issues.
- Optimize the performance, cost, and reliability of backend services with engineering teams.
- Lead incident response, post-mortem analysis, and long-term remediation efforts.
- Identify and eliminate operational toil while promoting self-service and operational maturity.
- Coordinate infrastructure roadmaps, security standards, and department-wide PCI and SOC compliance initiatives.
- Mentor engineers and collaborate with cross-functional teams.
Requirements
- 12+ years of software engineering experience with significant Site Reliability Engineering or DevOps experience.
- Expert experience with Google Cloud Platform services, including GKE, Compute Engine, Cloud Run, and Pub/Sub.
- Proficiency with infrastructure-as-code tools such as Terraform or Pulumi.
- Deep experience with Kubernetes, container orchestration, and service mesh architectures.
- Strong experience with monitoring and observability tools such as Datadog, Prometheus, Grafana, and Cloud Monitoring.
- Experience designing and managing high-throughput distributed systems.
- Strong problem-solving skills, communication skills, and ability to make high-stakes technical trade-offs.
- Demonstrated ability to mentor engineers.
- Preferred: experience in high-velocity customer-focused environments, functional programming with Clojure or ClojureScript, restaurant technology, hospitality, AI-driven SaaS, and cloud security or compliance best practices.
Benefits
- Salary, equity, comprehensive healthcare, paid time off, and other benefits.
- Available throughout Canada; the posting is tagged remote.
- Specific salary range provided by the recruiting team based on location and years of experience.
Tech Stack
Categories
Site Reliability