16 hours ago
Fort Lauderdale, FL, USASenior
Responsibilities
- Define and evolve SLIs, SLOs, and error budgets with product teams to guide reliability decisions.
- Design observability for a distributed serverless system across metrics, logs, and traces.
- Build incident response practices including on-call, escalation, blameless post-mortems, and reliability improvement loops.
- Develop reliability automation to detect and remediate issues before they affect clients.
- Improve CI/CD pipelines and deployment automation using GitHub Actions.
- Support resilient design, capacity planning, canary and blue-green deployments, and rollback automation.
- Contribute to disaster recovery design and testing for client-facing environments.
- Partner with Security and Compliance on regulatory and audit-ready operational practices.
- Define and implement operational readiness criteria for client go-live.
- Mentor colleagues and promote shared operational ownership.
Requirements
- At least 5 years of experience in Site Reliability Engineering, DevOps, or production operations for distributed cloud systems.
- Substantial hands-on AWS experience.
- Practical experience defining and using SLIs, SLOs, and error budgets.
- Strong incident response experience, including on-call, pressure triage, and post-mortems.
- A well-developed approach to observability for distributed systems.
- Solid automation and scripting ability; Rust, TypeScript, or Python experience is suitable.
- Strong collaboration and communication skills with a pragmatic approach to speed and stability.
- Preferred experience with AWS serverless and event-driven architecture, including idempotency, retries, dead-letter queues, and failure handling.
- Preferred experience taking platforms from pre-launch to production and defining operational readiness criteria.
- Preferred experience building AI agents to monitor platform health.
- Strong CI/CD experience, ideally with GitHub Actions, including secrets management, least-privilege permissions, and deployment controls.
- Infrastructure as Code experience with Terraform or OpenTofu.
- Experience with canary, blue-green, or rollback-based progressive deployments.
- Experience designing and testing disaster recovery for production SaaS.
- Experience operating across multiple cloud providers or designing for portability.
- Experience operating B2B SaaS in regulated environments and familiarity with SOC 2, PCI DSS, or GDPR as applied to operations.
- Familiarity with AI-assisted engineering tools.
- AWS certification, particularly DevOps Engineer - Professional, Solutions Architect - Professional, or Security - Specialty, is a bonus.
Benefits
- Hybrid and flexible working arrangements are available for most employees.
- The role requires regular travel to Switzerland.
- The company offers an inclusive and equal-opportunity work environment.
Tech Stack
Categories
Site Reliability
About Avaloq
Avaloq develops core banking and digital wealth management platforms and provides banking-operations outsourcing for financial institutions, delivered via cloud SaaS, BPaaS, or on-premises deployments. Its clients include private, retail and investment banks; more than 170 institutions use its systems. Founded in 1985 and headquartered in Zürich, Avaloq is a subsidiary of NEC Corporation.
