Base Salary
$155k - $235k/yr
Responsibilities
- Own the reliability and availability of colocation server fleets, on-premises lab clusters, cloud environments, and customer-facing platform services.
- Provision and operate servers from bare metal, including OS configuration, network setup, storage management, and hardware troubleshooting.
- Operate high-speed interconnect environments including InfiniBand, RoCE, and high-speed Ethernet.
- Manage capacity planning and hardware lifecycles for assigned infrastructure domains.
- Own Terraform and Ansible infrastructure as code and configuration management, ensuring provisioning and changes are performed through code.
- Build host lifecycle management, fleet health checks, auto-remediation workflows, self-service tooling, and networking automation.
- Design and maintain monitoring dashboards, alerting, and service-level indicators using Prometheus/Grafana and DataDog.
- Participate in on-call rotations, triage incidents from bare metal through the application layer, and implement permanent reliability improvements.
- Produce root-cause analysis reports for P0/P1 incidents with tracked action items.
- Operate customer-facing platform services and support their quality-of-service and uptime commitments.
- Maintain runbooks, architecture diagrams, configuration documentation, access procedures, and troubleshooting guides.
- Partner with DevOps and engineering teams to support CI/CD performance, developer experience, and infrastructure risk awareness.
Requirements
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience, plus 5+ years in SRE, infrastructure engineering, or systems administration.
- Strong Linux systems knowledge covering networking, storage, systemd, package management, kernel parameters, and performance diagnostics.
- Hands-on experience with colocation or on-premises server infrastructure, including physical hardware, rack networking, and bare-metal provisioning.
- Production infrastructure-as-code experience with Terraform and/or Ansible.
- Operational Kubernetes experience covering cluster troubleshooting, workload management, storage, and networking.
- Experience building dashboards and alert rules with Prometheus and Grafana or DataDog, with an understanding of signal quality.
- Production-quality automation scripting with Python and/or Bash.
- Incident response experience including structured triage, root-cause analysis reports, and follow-through on action items.
- Strongly preferred experience operating customer-facing infrastructure or platform services with external reliability expectations.
- Strongly preferred cloud infrastructure operations across AWS, Azure, or GCP, including hybrid cloud and on-premises environments.
- Strongly preferred HPC job scheduler experience with Slurm, LSF, or an equivalent.
- Strongly preferred knowledge of high-speed interconnect fabrics such as InfiniBand, RoCE, or NVLink.
- Strongly preferred Go programming for SRE tooling and experience with large-scale infrastructure automation or AIOps-driven operations.
Benefits
- 6-month contract with potential conversion to full-time.
Tech Stack
Categories
About d-Matrix
d-Matrix is an AI infrastructure company building the next generation of inference computing for the era of generative and agentic AI. Founded in 2019, d-Matrix is rethinking AI inference from the ground up with a full-stack approach spanning silicon, systems, networking, and software. Its flagship products, including the Corsair™ inference platform and JetStream™ inference fabric, are purpose-built to deliver high-performance, low-latency, and energy-efficient AI inference at datacenter scale. The company has raised nearly $500 million from a global syndicate of leading venture capital firms, sovereign wealth funds, and strategic investors, including Playground Global, Bullhound Capital, M12 (Microsoft’s Venture Fund), SK hynix, Temasek, Qatar Investment Authority, and Singapore’s EDBI. The company’s most recent financing valued d-Matrix at approximately $2 billion. As AI shifts from training to inference, the demands on AI infrastructure are changing. d-Matrix is building the infrastructure required to power the next generation of real-time AI at scale.
