
Site Reliability Engineer
Jump Trading28 days ago
Singapore, SingaporeSenior
Responsibilities
- Own production deployment, configuration, and release processes.
- Drive performance, reliability, operability, scalability, and uptime improvements.
- Build and maintain tooling for deployment, orchestration, monitoring, and system diagnostics.
- Define observability, SLI/SLOs, and performance metrics with product owners.
- Use metrics and capacity planning to support scalable and reliable systems.
- Troubleshoot and resolve complex production incidents across engineering teams.
- Lead incident response, root cause analysis, and post-mortems.
- Influence architecture and align practices with global SRE teams.
- Document procedures, mentor peers, and provide cross-training.
- Manage operational risk for production changes and participate in shared on-call duties.
Requirements
- A degree in Computer Science, a related field, or equivalent professional experience.
- At least 5+ years of relevant IT operations experience in areas such as DevOps, SRE, Linux Systems Engineering, or Network Engineering.
- Expert-level proficiency in C++.
- Strong knowledge of Linux, including networking and system configuration, kernel internals, scheduling, and performance tuning.
- Strong understanding of routing, multicast, LLDP, VLANs, and Ethernet.
- A rigorous, detail-oriented approach and strong ownership of operational outcomes.
- Ability to handle shared operational and periodic on-call responsibilities.
- Reliable and predictable availability.