
Senior Platform Systems Engineer
Bettermode4 months ago
Remote, Canada or Toronto, CanadaSenior
Responsibilities
- Diagnose and remediate platform issues across Kubernetes/EKS, AWS, Cloudflare, networking, observability, OLAP/data systems, security controls, and deployment architecture.
- Own reproducible, secure, recoverable Kubernetes platform patterns and Terraform/OpenTofu workflows, including promotion, drift control, and policy-aware infrastructure changes.
- Design availability-zone- and topology-aware improvements for Aurora PostgreSQL and other data-plane systems.
- Build workload-level cost and resource observability covering infrastructure spend, network transfer, CPU/memory usage, and cross-AZ behavior.
- Develop production platform components in Go, Rust, or TypeScript, including Kubernetes controllers, Terraform plugins, telemetry collectors, proxies, and CLIs.
- Implement SOC 2, OWASP, GDPR, IAM least-privilege, encryption, secrets-management, network-segmentation, auditability, and data-protection controls.
- Support OLAP infrastructure and the migration from Pinot to ClickHouse.
- Design canary, rollback, degraded-mode, backup/restore, failover, disaster-recovery, and incident-response patterns.
- Participate in an every-other-week production on-call rotation for P0 incident response and post-incident learning.
Requirements
- Deep production experience with Kubernetes/EKS, Terraform/OpenTofu, AWS, and Cloudflare, including secure deployments, environment promotion, drift control, and operational ownership.
- Strong software engineering experience with backend, infrastructure, or distributed systems in production, including concurrency, performance, failure modes, and operational correctness.
- Professional production experience with Go or Rust for platform components; TypeScript experience is valuable.
- Strong understanding of Linux, TCP/IP, HTTP, HTTP/2, gRPC, connection behavior under load, and service-to-service networking.
- Strong security and compliance knowledge covering SOC 2, OWASP, GDPR, IAM boundaries, encryption, secrets management, and audit trails.
- Practical experience with disaster recovery, backups, restore testing, RTO/RPO trade-offs, failover, degraded-mode operation, and incident playbooks.
- Familiarity with Aurora PostgreSQL, MongoDB Atlas, Pinot, ClickHouse, or comparable database and OLAP systems.
- Bonus qualifications include Kubernetes controllers or operators, service meshes, proxies, network observability, cost attribution, eBPF, VPC Flow Logs, MLOps, KServe, Kubeflow, MLflow, GPU workloads, infrastructure migration, policy-as-code, GitOps generators, CDK-style definitions, and internal developer platforms.
Benefits
- Full-time employment with remote or hybrid work in Canada; employees within 40 km of headquarters work from the downtown Toronto office three days per week on Monday, Tuesday, and Wednesday.
- Canadian health benefits including dental and vision coverage for employees and their families from the first day.
- Unlimited paid vacation, paid parental leave, and bereavement leave.
- Company-provided equipment or access to an interest-free Device Upgrade Policy.
- Monthly Technology and Appreciation stipend.
- Downtown Toronto office with a shuttle from Union Station, complimentary snacks and coffee, games, dedicated seating, and a flexible work environment.
- Location-based compensation with annual reviews.
Tech Stack
About Bettermode
Bettermode builds a customer community platform for businesses to host discussions, Q&A, events, and help centers, with customization, apps, embeds, and analytics. It sells its software as a SaaS subscription and supports building standalone or embedded community experiences. Founded in 2018 and headquartered in Toronto, the company cites customers such as IBM, ASUS, Monday.com, and Webflow.