4 hours ago
Bogotá, ColombiaStaff+
Responsibilities
- Own the architecture and evolution of core infrastructure as traffic, data volume, and workload complexity grow.
- Build capacity models, run load and stress tests, diagnose bottlenecks, and improve throughput, latency, reliability, and cost efficiency.
- Design resilient AWS account, IAM, networking, multi-AZ, and multi-region architectures.
- Build and operate Kubernetes platforms, including cluster lifecycle, workload isolation, autoscaling, upgrades, and deployment reliability.
- Scale and optimize Aurora RDS for MySQL and Postgres, including queries, indexes, replication, failover, migrations, and recovery.
- Implement reliability patterns such as service-level objectives, error budgets, failure isolation, backpressure, load shedding, and safe retries.
- Participate in on-call rotations and lead technical recovery during serious incidents and outages.
- Design, test, and document disaster recovery, backup, restore, and failover mechanisms.
- Build infrastructure-as-code, deployment automation, operational tooling, observability, and safe infrastructure migrations.
- Improve cloud cost efficiency and build AI-assisted tooling for incident investigation, runbooks, anomaly analysis, and toil reduction.
- Write architecture proposals, evaluate tradeoffs through prototypes and benchmarks, and document shared infrastructure behavior and failure modes.
Requirements
- Bachelor’s degree in Computer Science or a similar technical field is required.
- 12+ years of experience across infrastructure, platform, site reliability, software development, or related engineering disciplines.
- Deep production expertise with AWS, including compute, IAM, multi-account architectures, VPC design, and private connectivity.
- Deep production expertise with Kubernetes, including cluster lifecycle, scheduling, resource management, autoscaling, networking, and troubleshooting; EKS is strongly preferred.
- Deep expertise with RDS/Aurora using MySQL and/or Postgres, including performance, indexing, replication, high availability, failover, backup, and recovery.
- Demonstrated delivery of infrastructure scaling improvements with measurable gains in capacity, latency, reliability, or cost efficiency.
- Strong coding and automation skills with Golang, Python, or similar languages, plus Terraform or equivalent infrastructure-as-code experience.
- Strong systems fundamentals in Linux, networking, DNS, TLS, storage, concurrency, and distributed-system failure modes.
- Experience operating 24/7 high-availability platforms, responding to incidents, and implementing postmortem remediation.
- Experience with disaster recovery, observability, load testing, capacity planning, and safe CI/CD practices.
- Active use of AI tooling in engineering or operations with the ability to verify generated code, recommendations, and operational actions.
- Ability to solve ambiguous system-wide technical problems and collaborate across engineering disciplines.
- Preferred qualifications include fintech, payments, or banking experience; multi-region architectures; chaos and failure testing; Prometheus, Grafana, Loki, or Tempo; internal platform and self-service tooling; and AI-assisted incident investigation or operational automation.
Benefits
- Full-time remote position.
- Monthly gross compensation range of $12,500-$20,800 based on location and experience.
- Opportunity to work on infrastructure supporting fintech and payment workloads.
- Collaborative culture with engineering, data, and other technically diverse teams.
Tech Stack
Categories
About Sezzle
Sezzle builds a buy now, pay later platform that lets consumers split purchases into interest-free installments online and in stores, integrated with e-commerce and retail merchants. The company earns revenue from merchant fees and related consumer charges and provides underwriting and payment processing services. Founded in 2016 and headquartered in Minneapolis, Sezzle is a public company dual-listed on Nasdaq and the ASX.
