3 months ago
Toronto, CanadaSenior
Responsibilities
- Design, implement, and maintain highly available and scalable infrastructure for projects, products, and customers.
- Monitor and analyze system performance and resolve bottlenecks and reliability issues.
- Automate infrastructure deployment and configuration management.
- Improve system reliability, security, and efficiency through proactive monitoring, capacity planning, and performance tuning.
- Troubleshoot and resolve complex infrastructure and application issues in production and test environments.
- Collaborate with software engineering teams to design resilient, scalable, and secure systems.
- Participate in an on-call rotation and respond to production incidents.
- Document system configurations, troubleshooting procedures, and operational guidelines.
Requirements
- Proven experience as a Site Reliability Engineer or in a similar role.
- Strong understanding of networking, operating systems, and cloud infrastructure.
- Experience with Site Reliability Engineering, system design, and distributed computing.
- Experience with programming languages and SDK ecosystems including NodeJS, Java, Python, Ruby, and Go.
- Experience with containerization technologies such as Docker and Kubernetes.
- Knowledge of infrastructure-as-code tools such as Terraform and Pulumi.
- Familiarity with monitoring and logging tools including Prometheus, Grafana, and the ELK stack.
- Experience with lower-level implementation details of relational databases; distributed SQL database experience with Google Cloud Spanner or CockroachDB is a bonus.
- Experience working with Git and GitHub.
- Experience with continuous integration and deployment systems.
- Strong problem-solving and troubleshooting skills.
- Excellent communication and collaboration abilities.
- Experience with authorization systems is an extra qualification.
Benefits
- Fully remote work across the US, Canada, and Europe with a flexible schedule accommodating different time zones.
- Stock options at an early-stage startup.
- Comprehensive healthcare benefits for US-based employees and other insurance.
- Twice-yearly travel for team offsites focused on team bonding and collaboration.
Categories
DevOpsSite Reliability
About AuthZed
AuthZed, a leader in permissions systems as a service is on a mission to help every organization build fast and secure authorization that scales. As the creators of the open-source project SpiceDB, AuthZed has established a scalable and consistent system for storing and computing permissions data—use it to build fine-grained authorization services.
