1 day ago
Remote, PolandStaff+
Responsibilities
- Own the architecture and evolution of core infrastructure as traffic, data volume, and workload complexity grow.
- Build capacity models, conduct load and stress tests, and improve throughput, latency, saturation, reliability, and cost efficiency.
- Design resilient AWS account, IAM, networking, multi-AZ, and multi-region architectures.
- Operate and improve the Kubernetes platform, including cluster lifecycle, workload isolation, autoscaling, upgrades, and deployment reliability.
- Scale and optimize Aurora RDS for MySQL and Postgres, including queries, indexes, connections, replication, failover, schema changes, and migrations.
- Define service-level objectives and error budgets and implement failure isolation, backpressure, load shedding, and safe retries.
- Participate in on-call rotations and lead technical recovery from serious incidents and outages.
- Design, test, and document disaster recovery, backup, restore, and failover mechanisms.
- Build infrastructure-as-code, deployment automation, observability, migration, rollback, and cost-optimization capabilities.
- Build and evaluate AI-assisted operational tooling for investigation, runbooks, anomaly analysis, and toil reduction.
- Write architecture proposals, benchmark alternatives, review shared-infrastructure changes, and document system behavior and failure modes.
Requirements
- Bachelor’s degree in Computer Science or a similar technical field is required.
- 12+ years of experience across infrastructure, platform, site reliability, software development, or related engineering disciplines.
- Deep production expertise with AWS, including compute, IAM, multi-account architectures, VPC design, and private connectivity.
- Deep production expertise with Kubernetes, including cluster lifecycle, scheduling, resource management, autoscaling, networking, and troubleshooting; EKS is strongly preferred.
- Deep experience with relational databases at scale, specifically RDS/Aurora with MySQL and/or Postgres.
- Demonstrated personal delivery of infrastructure scaling improvements with measurable capacity, latency, reliability, or cost gains.
- Strong coding and automation skills using Golang, Python, or similar languages, plus Terraform or equivalent infrastructure-as-code experience.
- Strong systems fundamentals in Linux, networking, DNS, TLS, storage, concurrency, and distributed-system failure modes.
- Experience operating 24/7 high-availability platforms, responding to incidents, and implementing postmortem remediation.
- Experience with disaster recovery, observability, load testing, capacity planning, and safe CI/CD practices.
- Active use of AI tooling in engineering or operations with the ability to verify generated code, recommendations, and actions.
- Ability to resolve ambiguous, system-wide technical problems across engineering disciplines.
- Preferred experience includes fintech, payments, or banking; multi-region architectures; chaos engineering; failure testing; Prometheus, Grafana, Loki, or Tempo; internal platforms; self-service tooling; progressive delivery; and AI-assisted incident automation.
Benefits
- Full-time remote position.
- Monthly gross compensation range of $12,500–$20,800 USD based on location and experience.
- Opportunities to work on infrastructure supporting fintech and payment workloads at scale.
- Company culture includes employees with interests such as music, yoga, cycling, cooking, golf, dogs, and rock climbing.
Tech Stack
Categories
DevOpsSite Reliability
About Sezzle
Sezzle builds a buy now, pay later platform that lets consumers split purchases into interest-free installments online and in stores, integrated with e-commerce and retail merchants. The company earns revenue from merchant fees and related consumer charges and provides underwriting and payment processing services. Founded in 2016 and headquartered in Minneapolis, Sezzle is a public company dual-listed on Nasdaq and the ASX.
