6 hours ago
Remote, GermanyStaff+
Responsibilities
- Own the architecture and evolution of core infrastructure as traffic, data volume, and workload complexity increase.
- Build and optimize AWS infrastructure, Kubernetes platforms, and Aurora RDS MySQL and Postgres databases.
- Perform capacity planning, load and stress testing, bottleneck diagnosis, database tuning, autoscaling, and cost optimization.
- Improve reliability through service-level objectives, error budgets, failure isolation, backpressure, load shedding, and safe retry behavior.
- Participate in on-call rotations and lead technical recovery during major production incidents, including full outages.
- Design, test, and document disaster recovery, backup, restore, and failover mechanisms.
- Build infrastructure-as-code, deployment automation, observability, operational tooling, and safe infrastructure migrations.
- Develop and evaluate AI-assisted tools for incident investigation, runbooks, anomaly analysis, and toil reduction with appropriate controls.
- Write architecture proposals, evaluate technology tradeoffs, prototype and benchmark solutions, and review changes affecting shared infrastructure.
Requirements
- Bachelor’s degree in Computer Science or a similar technical field is required.
- 12+ years of experience across infrastructure, platform, site reliability, software development, or related engineering disciplines.
- Deep production expertise with AWS, including compute, IAM, multi-account architectures, networking, VPC design, and private connectivity.
- Deep production expertise with Kubernetes, including cluster lifecycle, scheduling, resource management, autoscaling, networking, and troubleshooting.
- Deep expertise with RDS/Aurora using MySQL and/or Postgres, including query performance, indexing, replication, high availability, failover, and backup recovery.
- Track record of personally delivering infrastructure scaling improvements with measurable gains in capacity, latency, reliability, or cost efficiency.
- Strong coding and automation skills with Golang, Python, or similar languages, plus Terraform or equivalent infrastructure-as-code experience.
- Strong systems fundamentals in Linux, networking, DNS, TLS, storage, concurrency, and distributed-system failure modes.
- Experience operating 24/7 high-availability platforms, responding to incidents, conducting postmortems, and implementing disaster recovery.
- Willingness and demonstrated ability to participate in on-call rotations and recover production systems under pressure.
- Experience with observability, load testing, capacity planning, and safe CI/CD practices.
- Active use of AI tooling in engineering or operations and the ability to verify generated code, recommendations, and operational actions.
- Ability to take ambiguous technical problems from investigation through production delivery and collaborate across engineering disciplines.
- Preferred qualifications include fintech, payments, or banking experience; multi-region architectures; chaos engineering; failure testing; Prometheus, Grafana, Loki, or Tempo; internal platform capabilities; self-service tooling; progressive delivery; and AI-assisted incident automation.
Benefits
- Full-time remote role.
- Monthly gross compensation range of $12,500-$20,800 based on location and experience.
- Opportunity to work on fintech infrastructure handling demanding reliability, security, and audit requirements.
- Collaborative culture with engineers, data enthusiasts, innovators, and employees pursuing varied personal interests.
Tech Stack
Categories
About Sezzle
Sezzle builds a buy now, pay later platform that lets consumers split purchases into interest-free installments online and in stores, integrated with e-commerce and retail merchants. The company earns revenue from merchant fees and related consumer charges and provides underwriting and payment processing services. Founded in 2016 and headquartered in Minneapolis, Sezzle is a public company dual-listed on Nasdaq and the ASX.
