
Research Engineer - Pre-training
Pluralis Research10 days ago
Sydney, AustraliaMid Level
Responsibilities
- Implement and optimize distributed model-parallel pretraining using data, pipeline, and tensor parallelism across heterogeneous GPUs.
- Reduce communication overhead while maintaining model convergence over low-bandwidth, high-latency networks.
- Build robust checkpointing, state synchronization, and recovery systems that handle node churn.
- Create instrumentation and monitoring for throughput, bottlenecks, and model quality across hundreds of devices.
Requirements
- Hands-on experience training models across many devices in PyTorch using FSDP, DeepSpeed, Megatron, or an equivalent implementation.
- Understanding of data, tensor, and pipeline parallelism.
- Strong production-quality Python engineering skills, including concurrency, failure handling, and profiling.
- Evidence of execution through shipped systems, research code, open-source work, or serious personal projects.
- Professional-level written and spoken English proficiency.
- Experience training or serving large language models such as Nemotron, Qwen, or OLMo is preferred.
- Experience with P2P networking and NAT traversal is preferred.
- Experience with post-training and reinforcement learning is preferred.
- Experience with inference and serving systems is preferred.
- Experience at proprietary, open-weight, or open-source AI labs is preferred.
Benefits
- Remote-first work environment with globally distributed team members.
- Full visa sponsorship and relocation support to Australia or the United States may be available.
- Equity-heavy package with significant ownership for key technical contributors in addition to a high base salary.
- Applicants should be comfortable working across time zones.
Categories
About Pluralis Research
Pluralis Research builds a protocol and tooling for collaborative, decentralized training and ownership of foundation models, enabling training and serving across heterogeneous, consumer-grade devices over the internet. The team demonstrated Agora, a permissionless run that pretrained an 8B model from scratch without any single participant holding the full weights. The privately held company is backed by Union Square Ventures and focuses on developer-facing infrastructure for large-scale model training and serving.