Responsibilities
- Design, implement, and maintain production GPU system software across kernel drivers, runtimes, firmware interfaces, host management, debugging, tracing, profiling, and performance measurement.
- Own accelerator features from proof of concept and architecture through implementation, validation, bring-up, qualification, and deployment.
- Debug complex failures involving GPUs, CPUs, memory, PCIe, interconnects, IOMMU, firmware, operating systems, runtimes, and distributed workloads.
- Profile workloads and optimize initialization, memory movement, scheduling, synchronization, communication, recovery, and device utilization.
- Build automated tests, telemetry, dashboards, health checks, and diagnostic tools to expose failures and regressions.
- Collaborate with silicon, firmware, compiler, library, machine-learning framework, platform, validation, and production engineering teams.
- Improve resilience through error detection, isolation, retry, reset, repair, graceful degradation, and operational procedures.
- Review low-level designs and code, document hardware/software contracts, and share performance and debugging practices.
Requirements
- Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
- At least 3 years of professional programming experience in C or C++ for low-level, embedded, kernel, driver, runtime, or performance-critical software.
- Strong understanding of computer architecture, operating systems, concurrency, memory hierarchies, DMA, interrupts, and device I/O.
- Hands-on experience with GPU or accelerator software in areas such as kernel drivers, runtimes, firmware, libraries, collective communication, or performance tooling.
- Ability to diagnose system failures using traces, logs, profilers, debuggers, counters, and controlled experiments.
- Experience delivering and maintaining production-quality features across hardware and software teams.
- Preferred experience with CUDA, ROCm, Level Zero, OpenCL, Triton, CUTLASS, or comparable accelerator programming and runtime environments.
- Preferred knowledge of GPU scheduling, memory management, virtualization, PCIe, coherent interconnects, NUMA, or multi-GPU topology.
- Preferred experience with pre-silicon development, board bring-up, firmware communication, baseboard management controllers, or fleet qualification.
- Familiarity with distributed training or inference and topology-aware communication libraries is preferred.
About ByteDance
ByteDance is a global incubator of platforms at the cutting edge of commerce, content, entertainment and enterprise services - over 2.5bn people interact with ByteDance products including TikTok. Creation is the core of ByteDance's purpose. Our products are built to help imaginations thrive. This is doubly true of the teams that make our innovations possible. Together, we inspire creativity and enrich life - a mission we aim towards achieving every day. At ByteDance, we create together and grow together. That's how we drive impact - for ourselves, our company, and the users we serve. We are committed to building a safe, healthy and positive online environment for all our users. We have over 110,000 employees based in more than 30 countries globally. Join us.
