
Site Reliability Engineer
CXM Direct LLC2 months ago
Remote, AmericasMid Level
Responsibilities
- Participate in the on-call rotation and lead incident response for production trading systems.
- Investigate incidents, perform root cause analysis, and implement preventive actions.
- Build and maintain Grafana dashboards, Prometheus alerts, and operational health views.
- Instrument .NET services to improve telemetry, metrics, logging, and service-health visibility.
- Define and monitor SLIs, SLOs, and error budgets.
- Troubleshoot .NET/C# applications, Windows Server, Aurora PostgreSQL, AWS infrastructure, and deployments.
- Improve deployment safety, release automation, rollback strategies, and operational automation.
- Partner with software engineers to improve application operability, resilience, and fault isolation.
- Create and maintain runbooks, operational documentation, and incident-response procedures.
Requirements
- 3–5 years of experience in site reliability, production engineering, or a related area.
- Strong experience debugging and supporting .NET/C# applications in production.
- Hands-on experience with Windows Server environments.
- Strong PowerShell scripting skills and experience with Python or Bash.
- Experience with Grafana, Prometheus, Loki, or equivalent monitoring and observability tools.
- Understanding of metrics, logging, tracing, and alerting best practices.
- Experience with modern CI/CD pipelines, deployment strategies, release automation, and rollback mechanisms.
- Experience working with AWS and hands-on experience with Terraform or other infrastructure-as-code tools.
- Experience troubleshooting Aurora PostgreSQL or other relational database platforms.
- Practical experience with SLIs, SLOs, error budgets, incident response, root cause analysis, alert design, and production operations.
- Preferred: experience supporting high-availability or low-latency financial or trading systems.
- Preferred: familiarity with MetaTrader or financial technology platforms.
- Preferred: experience with distributed systems and microservices.
- Preferred: knowledge of OpenTelemetry or similar observability frameworks.
- Preferred: exposure to Docker, Kubernetes, or containerized environments.
Benefits
- Remote position for the Americas, with LatAm preferred.
- Working hours aligned with Americas time zones from UTC-3 to UTC-8.
- On-call rotation aligned with the London trading day.
- Full-time, permanent employment.
- Opportunity to work on mission-critical trading infrastructure and influence reliability strategy and engineering practices.
Tech Stack
Categories
Site Reliability