4 hours ago
Remote, MexicoStaff+
Responsibilities
- Own the technical architecture and evolution of core infrastructure as traffic, data volume, and workload complexity grow.
- Build capacity models, run load and stress tests, diagnose bottlenecks, and improve throughput, latency, reliability, and cost efficiency.
- Design resilient AWS account, IAM, networking, multi-AZ, and multi-region architectures.
- Build and operate Kubernetes platforms, including cluster lifecycle, workload isolation, resource allocation, autoscaling, upgrades, and deployment reliability.
- Scale and optimize Aurora RDS for MySQL and Postgres through query and index tuning, connection and replication improvements, failover planning, and safe migrations.
- Implement reliability engineering practices including service-level objectives, error budgets, failure isolation, backpressure, load shedding, and safe retries.
- Participate in on-call rotations, recover systems during major incidents, and implement corrective actions from postmortems.
- Design, test, and document disaster recovery, backup, restore, and failover mechanisms against recovery objectives.
- Build infrastructure-as-code, deployment, provisioning, upgrade, and operational automation to reduce manual toil.
- Improve observability through metrics, logs, traces, dashboards, alerts, and visibility into customer impact.
- Deliver safe infrastructure migrations with phased rollouts, validation, compatibility checks, and rollback paths.
- Build and evaluate AI-assisted tooling for incident investigation, capacity analysis, runbook automation, anomaly analysis, and toil reduction.
- Write architecture proposals, prototype and benchmark technical solutions, review shared-infrastructure changes, and document system behavior and failure modes.
Requirements
- Bachelor's degree in Computer Science or a similar technical field is required.
- 12+ years of experience across infrastructure, platform, site reliability, software development, or related engineering disciplines.
- Deep production expertise with AWS, including compute, IAM, multi-account architectures, networking, VPC design, and private connectivity.
- Deep production expertise with Kubernetes, including cluster lifecycle, scheduling, resource management, autoscaling, networking, and troubleshooting; EKS is strongly preferred.
- Deep experience with RDS/Aurora using MySQL and/or Postgres, including query performance, indexing, connection management, replication, high availability, failover, backup, and recovery.
- Demonstrated personal delivery of infrastructure scaling improvements with measurable gains in capacity, latency, reliability, or cost efficiency.
- Strong coding and automation skills using Golang, Python, or similar languages, plus infrastructure-as-code experience with Terraform or an equivalent.
- Strong systems fundamentals covering Linux, networking, DNS, TLS, storage, concurrency, and distributed-system failure modes.
- Experience operating 24/7 high-availability platforms with hands-on incident response and postmortem remediation.
- Willingness and ability to participate in on-call rotations and recover production systems under pressure.
- Experience implementing and testing disaster recovery, including data restoration and service-recovery validation.
- Practical experience with observability, load testing, capacity planning, and safe CI/CD practices for shared infrastructure.
- Active use of AI tooling in engineering or operations, including verification of generated code, recommendations, and operational actions.
- Ability to take ambiguous, system-wide technical problems from investigation through production delivery and collaborate across engineering disciplines.
- Preferred qualifications include fintech, payments, or banking experience; multi-region architectures; chaos engineering; failure testing; distributed data recovery tradeoffs; Prometheus, Grafana, Loki, or Tempo; internal platforms; self-service tooling; progressive delivery; and AI-assisted operational automation.
Benefits
- Remote, full-time role.
- Sezzle offers a collaborative culture centered on purpose-driven employees, high standards, innovation, accountability, and measurable results.
- The company highlights employee interests and activities including music, yoga, cycling, cooking, golf, dogs, and rock climbing.
Tech Stack
Categories
About Sezzle
Sezzle builds a buy now, pay later platform that lets consumers split purchases into interest-free installments online and in stores, integrated with e-commerce and retail merchants. The company earns revenue from merchant fees and related consumer charges and provides underwriting and payment processing services. Founded in 2016 and headquartered in Minneapolis, Sezzle is a public company dual-listed on Nasdaq and the ASX.
