20 hours ago
Base Salary
$175k - $265k/yr
Responsibilities
- Own reliability and availability across colocation servers, on-premises lab clusters, cloud environments, and customer-facing platform services.
- Provision and configure servers, operating systems, networking, storage, hardware, and auto-scaling Kubernetes environments.
- Lead capacity planning, hardware lifecycle management, cloud-spend tracking, and workload-placement decisions.
- Automate provisioning, deployment, and operational changes using Terraform and/or Ansible, including shared infrastructure-as-code modules.
- Build automation for host lifecycle management, fleet health checks, auto-remediation, self-service tooling, and networking automation.
- Design and maintain monitoring, alerting, dashboards, and service-level indicators using observability tools.
- Participate in on-call incident response, resolve issues from bare metal through the application layer, and produce root-cause analyses for major incidents.
- Support internal and external platform services, uptime commitments, operational runbooks, and customer deployments.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- At least 7 years of experience in SRE, infrastructure engineering, or systems administration.
- Strong Linux systems experience with colocation or on-premises server infrastructure, networking, storage, systemd, kernel parameters, performance diagnostics, physical hardware, rack networking, and bare-metal provisioning.
- Production experience writing and maintaining Terraform and/or Ansible configurations.
- Operational Kubernetes experience including cluster troubleshooting, workload management, storage, and networking.
- Experience with Prometheus/Grafana, Datadog, Splunk, or equivalent observability tooling, including dashboard and alert-rule creation.
- Production-quality Python and/or Bash scripting combined with incident-response, triage, root-cause analysis, and remediation experience.
- Preferred qualifications include customer-facing infrastructure operations, hybrid cloud and on-premises environments, production AIOps or intelligent alerting, HPC schedulers such as Slurm or LSF, high-speed interconnects such as InfiniBand, RoCE, or NVLink, and large-scale infrastructure automation.
Tech Stack
AnsibleAWSAzureBashDatadogGoogle Cloud PlatformGrafanaKubernetesLinuxPrometheusPythonSplunkTerraform
Categories
DevOpsSite Reliability
About d-Matrix
d-Matrix builds AI inference computing platforms for data centers, combining custom silicon with systems, networking, and software. Its flagship Corsair platform and JetStream fabric focus on low-latency, energy-efficient generative AI inference at scale. Founded in 2019 and headquartered in Santa Clara, California, the privately held company sells hardware with accompanying software to cloud providers and enterprises deploying large AI models.
