Sezzle

Principal Infrastructure Engineer

Sezzle
Apply
4 hours ago
Remote, BrazilStaff+

Responsibilities

  • Own the architecture and evolution of core infrastructure as traffic, data volume, and workload complexity grow.
  • Build capacity models, conduct load and stress testing, and diagnose performance bottlenecks across compute, networking, Kubernetes, and databases.
  • Design resilient AWS account, IAM, networking, multi-AZ, and multi-region architectures.
  • Build and operate Kubernetes platforms, including lifecycle automation, workload isolation, resource allocation, autoscaling, upgrades, and deployment reliability.
  • Scale and optimize Aurora RDS for MySQL and Postgres through query and index tuning, connection and replication improvements, capacity planning, failover improvements, and safe migrations.
  • Improve reliability with service-level objectives, error budgets, failure isolation, backpressure, load shedding, and safe retry behavior.
  • Participate in on-call rotations and lead technical recovery during major incidents and outages.
  • Design, test, and document disaster recovery, backup, restore, and failover mechanisms against recovery objectives.
  • Build infrastructure-as-code and operational automation for provisioning, configuration, deployment, upgrades, and recovery.
  • Improve observability through metrics, logs, traces, dashboards, and actionable alerts.
  • Deliver safe infrastructure migrations with phased rollouts, validation, compatibility checks, and rollback paths.
  • Improve cloud cost efficiency through resource sizing, utilization, autoscaling, and storage optimization.
  • Build and evaluate AI-assisted tooling for incident investigation, runbooks, anomaly analysis, and toil reduction with appropriate access controls and auditability.
  • Write architecture proposals, evaluate technology tradeoffs, prototype and benchmark solutions, review infrastructure changes, and document system operations and failure modes.

Requirements

  • Bachelor’s degree in Computer Science or a similar technical field is required.
  • At least 12 years of experience across infrastructure, platform, site reliability, software development, or related engineering disciplines is required.
  • Deep production expertise with AWS, including compute, IAM, multi-account architectures, networking, VPC design, and private connectivity.
  • Deep production expertise with Kubernetes, including cluster lifecycle, scheduling, resource management, autoscaling, networking, and troubleshooting; EKS is strongly preferred.
  • Deep experience with RDS/Aurora MySQL and/or Postgres at scale, including query performance, indexing, connection management, replication, high availability, failover, backup, and recovery.
  • Demonstrated personal delivery of infrastructure scaling improvements with measurable gains in capacity, latency, reliability, or cost efficiency.
  • Strong coding and automation skills using Golang, Python, or similar languages, plus infrastructure-as-code experience with Terraform or an equivalent.
  • Strong systems fundamentals in Linux, networking, DNS, TLS, storage, concurrency, and distributed-system failure modes.
  • Experience operating a 24/7 high-availability platform with hands-on incident response and postmortem remediation.
  • Willingness and ability to participate in on-call rotations and recover production systems under pressure.
  • Experience implementing and testing disaster recovery against defined recovery objectives.
  • Practical experience with observability, load testing, capacity planning, and safe CI/CD practices for shared production infrastructure.
  • Active use of AI tooling in engineering or operations and the judgment to verify generated code, recommendations, and operational actions.
  • Ability to take ambiguous technical problems from investigation through production delivery and collaborate across engineering disciplines.
  • Preferred qualifications include fintech, payments, or banking experience; multi-region architecture, chaos engineering, and failure testing; Prometheus, Grafana, Loki, or Tempo; internal platform and self-service tooling; progressive delivery; and AI-assisted operational automation.

Benefits

  • Full-time remote role.
  • Monthly gross compensation range of $12,500-$20,800, based on location and experience.
  • Opportunity to work on infrastructure supporting fintech and payments workloads with significant reliability, performance, and scale requirements.

Tech Stack

Categories

Sezzle

About Sezzle

201-500 employees

Sezzle builds a buy now, pay later platform that lets consumers split purchases into interest-free installments online and in stores, integrated with e-commerce and retail merchants. The company earns revenue from merchant fees and related consumer charges and provides underwriting and payment processing services. Founded in 2016 and headquartered in Minneapolis, Sezzle is a public company dual-listed on Nasdaq and the ASX.

Contact me