CXM Direct LLC

Site Reliability Engineer

CXM Direct LLC
Apply
2 months ago
Remote, AmericasMid Level

Responsibilities

  • Participate in the on-call rotation and lead incident response for production trading systems.
  • Investigate incidents, perform root cause analysis, and implement preventive actions.
  • Build and maintain Grafana dashboards, Prometheus alerts, and operational health views.
  • Instrument .NET services to improve telemetry, metrics, logging, and service-health visibility.
  • Define and monitor SLIs, SLOs, and error budgets.
  • Troubleshoot .NET/C# applications, Windows Server, Aurora PostgreSQL, AWS infrastructure, and deployments.
  • Improve deployment safety, release automation, rollback strategies, and operational automation.
  • Partner with software engineers to improve application operability, resilience, and fault isolation.
  • Create and maintain runbooks, operational documentation, and incident-response procedures.

Requirements

  • 3–5 years of experience in site reliability, production engineering, or a related area.
  • Strong experience debugging and supporting .NET/C# applications in production.
  • Hands-on experience with Windows Server environments.
  • Strong PowerShell scripting skills and experience with Python or Bash.
  • Experience with Grafana, Prometheus, Loki, or equivalent monitoring and observability tools.
  • Understanding of metrics, logging, tracing, and alerting best practices.
  • Experience with modern CI/CD pipelines, deployment strategies, release automation, and rollback mechanisms.
  • Experience working with AWS and hands-on experience with Terraform or other infrastructure-as-code tools.
  • Experience troubleshooting Aurora PostgreSQL or other relational database platforms.
  • Practical experience with SLIs, SLOs, error budgets, incident response, root cause analysis, alert design, and production operations.
  • Preferred: experience supporting high-availability or low-latency financial or trading systems.
  • Preferred: familiarity with MetaTrader or financial technology platforms.
  • Preferred: experience with distributed systems and microservices.
  • Preferred: knowledge of OpenTelemetry or similar observability frameworks.
  • Preferred: exposure to Docker, Kubernetes, or containerized environments.

Benefits

  • Remote position for the Americas, with LatAm preferred.
  • Working hours aligned with Americas time zones from UTC-3 to UTC-8.
  • On-call rotation aligned with the London trading day.
  • Full-time, permanent employment.
  • Opportunity to work on mission-critical trading infrastructure and influence reliability strategy and engineering practices.

Tech Stack

AWSBashC#DockerGrafanaKubernetes.NETPowerShellPrometheusPythonTerraformWindows

Categories

Site Reliability
CXM Direct LLC

About CXM Direct LLC

201-500 employees
Contact me