
HPC Data Center Developer
Jump Trading28 days ago
Chicago, IL, USA or New York, NY, USASenior
Base Salary
$150k - $200k/yr
Responsibilities
- Design, develop, and maintain automation for onboarding servers, network switches, rack PDUs, CDUs, and environmental sensors.
- Build end-to-end hardware provisioning workflows covering discovery, configuration, validation, and production readiness.
- Develop power and cooling capacity-planning tools and outage-simulation tools for HPC facilities.
- Maintain operational tooling for hardware lifecycle tracking, inventory, spares, change management, and diagnostics.
- Integrate telemetry from data center infrastructure and colocation providers into centralized monitoring and metrics systems.
- Implement production-grade monitoring and alerting instrumentation with the Operations Lead.
- Partner with HPC Planning, Engineering, and Operations teams on automation and monitoring requirements.
- Integrate data center automation with compute, storage, and network provisioning systems.
- Own system reliability, respond to failures, improve tooling based on operational feedback, and maintain comprehensive documentation.
- Use AI tools daily for development, code review, data analysis, debugging, documentation, and data center operations applications.
- Participate in coordinated maintenance operations, including evenings and weekends.
Requirements
- At least 5 years of professional experience in production engineering, infrastructure automation, or site reliability engineering, preferably in HPC or large-scale data center environments.
- Proven experience building and shipping maintained, reliable production automation and tooling.
- Experience automating hardware provisioning and lifecycle management for servers, network devices, and power or cooling infrastructure.
- Strong understanding of data center power distribution, air and liquid cooling, environmental monitoring, and structured cabling.
- Experience with IPMI, BMC, Redfish, SNMP, and vendor APIs for hardware discovery, configuration, and telemetry collection.
- High proficiency in Golang and at least one additional programming language such as Python.
- Strong Linux systems knowledge, including administration, networking, storage, process management, log analysis, and OS-level troubleshooting.
- Experience with Grafana and observability platforms such as Prometheus or InfluxDB, including custom integrations and exporters.
- Experience with configuration management and infrastructure-as-code tools such as SaltStack, Ansible, or Terraform.
- Understanding of L2/L3 networking, VLANs, BGP, SNMP, and switch or router configuration, including Arista and Cisco equipment.
- Experience consuming vendor APIs, normalizing heterogeneous data sources, and building data pipelines for metrics and reporting.
- Experience with ClickHouse and MySQL, including queries, schema design, and database tooling.
- Experience using GitHub for version control, code review, CI/CD workflows, and collaborative development.
- Demonstrated professional use of AI tools, including LLM-based coding assistants or AI-driven analytics.
- Strong root-cause-analysis ability, communication skills, work quality standards, and ability to work under pressure.
- Reliable availability for evenings and weekends when required.
- Bachelor's degree preferred.
Benefits
- Discretionary bonus eligibility.
- Medical, dental, and vision insurance.
- HSA, FSA, and dependent care options.
- Employer-paid group term life and AD&D insurance, plus voluntary life and AD&D insurance.
- Paid vacation and holidays.
- Retirement plan with employer match.
- Paid parental leave.
- Wellness programs.
- On-site work five days per week in Chicago or New York.
- Annual base salary range of $150,000 to $200,000 USD.