
Staff Platform Engineer
Checkfront10 hours ago
Remote, CanadaStaff+
Responsibilities
- Own and evolve AWS infrastructure, networking, compute, data services, observability, CI/CD, and operational tooling.
- Define and maintain infrastructure as code using Pulumi and TypeScript across the platform.
- Support containerized application platforms, deployment pipelines, rollback mechanisms, and runtime configuration.
- Operate PostgreSQL data infrastructure, including connection pooling, backups, disaster recovery, replication, and safe schema changes.
- Improve release reliability through database migrations, deployment workflows, environment management, and production readiness checks.
- Lead observability and incident-readiness efforts involving alerting, dashboards, SLOs, runbooks, incident response, and post-incident follow-up.
- Improve platform security, cost efficiency, maintainability, scalability, and developer experience.
- Mentor engineers on infrastructure, operations, reliability, and production ownership.
Requirements
- Deep production experience with AWS and services including ECS/Fargate, RDS/Aurora PostgreSQL, VPC networking, load balancing, IAM, KMS, Secrets Manager, CloudFront, and WAF.
- Experience designing and operating globally distributed systems with multi-region availability and disaster recovery procedures.
- Strong infrastructure-as-code experience; Pulumi with TypeScript is ideal, while Terraform or another mature IaC approach is also valuable.
- Strong operational knowledge of PostgreSQL, including performance investigation, connection pooling, backups, replication, locking, migrations, and safe schema-change practices.
- Experience designing and maintaining CI/CD systems, preferably with GitHub Actions, OIDC-based cloud authentication, container builds, environment promotion, required checks, and deployment gates.
- Experience supporting containerized production workloads and improving deployment safety, rollback strategies, and runtime reliability.
- Strong observability and incident response experience involving metrics, logs, traces, alerting, dashboards, runbooks, and post-incident learning.
- Ability to work effectively in ambiguity, make pragmatic tradeoffs, and communicate clearly with infrastructure specialists and product engineers.
- Track record of raising engineering standards through reusable patterns, documentation, automation, mentoring, and technical leadership.
Benefits
- Opportunity to shape the infrastructure and operating standards for a new product at an early stage.
- High-autonomy environment with meaningful influence over how Manifest operates, deploys, scales, and responds to incidents.