
Research Engineer - Post-Training
Pluralis Research3 hours ago
Sydney, AustraliaSenior
Responsibilities
- Build the end-to-end RL post-training stack, including geo-distributed rollout ingestion, reward computation, policy updates, and distributing updated weights across the network.
- Adapt RL algorithms to asynchronous, high-latency, partially trusted generation through staleness tolerance, off-policy corrections, and communication-efficient policy updates.
- Build evaluations to measure model improvement and ship the first decentralized post-trained model as a public artifact.
- Set technical direction and drive implementation of the post-training system and algorithms.
Requirements
- Hands-on experience running RL post-training on large language models using RLHF, RLVR, or reasoning-focused RL.
- Direct systems experience with rollout generation, asynchronous training loops, and weight synchronization.
- Production-quality Python and PyTorch engineering experience, including concurrency, failure handling, and profiling.
- Research ability demonstrated through publications or detailed, defensible unpublished work in RL post-training, asynchronous or distributed RL, or nearby fields.
- Professional-level written and spoken English proficiency.
- Ability to work comfortably across global time zones.
Benefits
- Equity-heavy package with significant ownership for key technical contributors in addition to a high base salary.
- Remote-first work environment with globally distributed team members.
- Optional full visa sponsorship and relocation support to Australia or the United States.
- Opportunity to work on largely unpublished problems in decentralized training and serving of frontier models.
- The role is remote across the world, with main teams in Australia and North America.
Categories
AI ResearchML Engineering
About Pluralis Research
Pluralis is developing a protocol that facilitates collaborative training and ownership of foundation models.