
Site Reliability Engineer
Jump Trading28 days ago
Singapore, SingaporeSenior
Responsibilities
- Own the production environment and drive performance, reliability, and operability improvements.
- Monitor and troubleshoot large-scale trading systems and exchange connectivity.
- Build and maintain DevOps tooling for configuration management, process management, deployment, monitoring, data collection, and analysis.
- Use firm-wide metrics to improve system scalability and performance.
- Collaborate across the technology organization to investigate and resolve complex system problems.
- Coordinate production changes and manage incidents with Risk Management and Operational Trading Support.
- Interact directly with traders to communicate technology changes, manage incidents, and troubleshoot problems.
- Work with the Clearing team to reconcile trades and position breaks.
- Assess and manage operational risk for production changes.
- Define and document processes and procedures.
- Mentor and cross-train other technical operations SREs.
- Participate in shared operational and periodic on-call duties.
Requirements
- Degree in Computer Science, a related field, or equivalent professional experience.
- At least 5 years of relevant IT operations experience in areas such as DevOps, SRE, Linux Systems Engineering, or Network Engineering.
- At least 3 years of experience with Python and shell scripting.
- Strong understanding of the Linux operating system, including networking and system configuration, kernel internals, scheduling, and performance tuning.
- Strong understanding of networking concepts including routing, multicast, LLDP, VLAN tagging, and Ethernet.
- Familiarity with C++ is helpful but not required.
- Detail-oriented and rigorous approach to operations, with a strong sense of ownership and urgency.
- Ability to maintain reliable availability and participate in shared operational and periodic on-call duties.