2 months ago
Sydney, AustraliaStaff+
Responsibilities
- Lead production support and site reliability engineering for mission-critical digital asset platforms and transaction-processing systems.
- Own incident management from detection through resolution and post-incident review, reducing business impact and recurrence.
- Analyze platform logic, diagnose root causes, and deliver scalable system enhancements.
- Partner with product owners, business analysts, and engineering teams to document impacts, define support requirements, and align priorities.
- Evaluate risk exposures, security vulnerabilities, and platform gaps and propose mitigation strategies.
- Develop automation scripts and tooling using Python, Java, or equivalent languages to reduce operational toil and improve observability.
- Mentor junior engineers, guide workload allocation, and raise the engineering standards of the support function.
Requirements
- 6–10 years of hands-on experience in production support, SRE, or platform engineering supporting high-availability, mission-critical systems at scale.
- Deep expertise in incident management and systems-level troubleshooting, including Linux diagnostics and real-time or distributed-system debugging.
- Practical production experience with monitoring and observability platforms such as Splunk, Grafana, ELK, Prometheus, OTEL, or GCO.
- Scripting and automation capability in Python, Java, or a comparable language, with examples of reducing operational toil or accelerating incident response.
- Strong SQL skills for data investigation, platform diagnostics, and incident analysis across complex datasets.
- Experience leading requirements gathering and system-scoping efforts across multiple concurrent large-scale technology programs.
- Ability to communicate platform issues, business impact, and recommendations clearly to technical and business audiences.
- Preferred experience with markets technology, electronic trading, payments platforms, FinTech exchanges, settlement flows, real-time financial data pipelines, blockchain-connected infrastructure, digital wallets, distributed ledger technologies, SLOs, error budgets, or reliability frameworks.
Benefits
- Hybrid working arrangement with three days in the office and two days working remotely.
- Technical ownership and strategic input into the reliability, architecture, and evolution of digital asset platforms.
- Exposure to blockchain, digital assets, and high-availability distributed systems at global scale.
- Structured learning and development programs supporting technical growth and career progression within technology leadership.
- Comprehensive benefits package including financial wellbeing provisions, wellbeing support, and family-oriented benefits.
- Full-time position with mentoring and cross-functional collaboration opportunities.
About Citi
Citi's mission is to serve as a trusted partner to our clients by responsibly providing financial services that enable growth and economic progress. Our core activities are safeguarding assets, lending money, making payments and accessing the capital markets on behalf of our clients. We have over 200 years of experience helping our clients meet the world's toughest challenges and embrace its greatest opportunities. We are Citi, the global bank – an institution connecting millions of people across hundreds of countries and cities. For information on Citi’s commitment to privacy, visit on.citi/privacy.
