
Research Engineer - Pre-training
Pluralis Research3 hours ago
Sydney, AustraliaMid Level
Responsibilities
- Implement and optimize distributed model-parallel pretraining using data, pipeline, and tensor parallelism across heterogeneous GPUs.
- Reduce communication overhead while maintaining model convergence over low-bandwidth, high-latency networks.
- Build robust checkpointing, state synchronization, and recovery systems that handle node churn.
- Create instrumentation and monitoring for throughput, bottlenecks, and model quality across hundreds of devices.
Requirements
- Hands-on experience training models across many devices in PyTorch using FSDP, DeepSpeed, Megatron, or an equivalent implementation.
- Understanding of data, tensor, and pipeline parallelism.
- Strong production-quality Python engineering skills, including concurrency, failure handling, and profiling.
- Evidence of execution through shipped systems, research code, open-source work, or serious personal projects.
- Professional-level written and spoken English proficiency.
- Experience training or serving large language models such as Nemotron, Qwen, or OLMo is preferred.
- Experience with P2P networking and NAT traversal is preferred.
- Experience with post-training and reinforcement learning is preferred.
- Experience with inference and serving systems is preferred.
- Experience at proprietary, open-weight, or open-source AI labs is preferred.
Benefits
- Remote-first work environment with globally distributed team members.
- Full visa sponsorship and relocation support to Australia or the United States may be available.
- Equity-heavy package with significant ownership for key technical contributors in addition to a high base salary.
- Applicants should be comfortable working across time zones.
Categories
About Pluralis Research
Pluralis is developing a protocol that facilitates collaborative training and ownership of foundation models.