1 month ago
Sydney, AustraliaSenior
Responsibilities
- Own the model inference stack end to end for in-house and open-source models.
- Build high-throughput inference servers serving millions of generations per day at under 200ms latency.
- Optimize GPU utilization through batching, quantization, and custom CUDA kernels.
- Test and productionize LoRAs and in-house models for tens of millions of users.
- Drive work from user understanding and idea development through implementation and iteration.
Requirements
- 5+ years of experience building software at scale, focused on ML inference or GPU-accelerated systems.
- Deep familiarity with GPU inference, including batching, quantization, and serving frameworks such as vLLM, TensorRT, or Triton.
- Ability to own projects end to end and deliver with urgency and high agency.
Benefits
- Top-of-market compensation with meaningful equity upside; the listed range is cash plus equity with superannuation on top.
- In-person work in Sydney, Australia, with visa sponsorship and relocation assistance for people moving to Sydney.
- Real ownership from day one and opportunities for scope, responsibility, and compensation growth.
- Company card for food, coffee, tools, and other work-related needs.
- Daily team lunch and dinner at the office.
- Unlimited workspace budget.
- Three-day paid work trial in Sydney, with travel and accommodation covered if needed.
