7 hours ago
Singapore, SingaporeSenior
Responsibilities
- Design and operate distributed training infrastructure for neural operator architectures on NVIDIA DGX B200 systems.
- Optimize training pipelines for throughput, fault tolerance, checkpointing, gradient accumulation, multi-node synchronization, and cost efficiency.
- Build experiment tracking and observability systems for training runs, hyperparameter sweeps, and model performance.
- Optimize data loading and cloud-storage I/O for large-scale mesh datasets through prefetching, caching, and format optimization.
- Build model-serving and packaging infrastructure for pre-trained Large Physics Models, including customer deployment and fine-tuning capabilities.
- Ensure reproducible model checkpoints and reliable deployment behavior across customer environments.
- Improve Research developer experience through debugging tools, fast iteration cycles, and collaboration on shared infrastructure standards.
Requirements
- At least 5 years of experience building and operating ML infrastructure at scale.
- Deep expertise in distributed training, including NCCL debugging, collective communication, FSDP, DDP, and pipeline parallelism.
- Strong systems fundamentals across Linux, networking, NVLink, InfiniBand, storage I/O, profiling, and performance optimization.
- Production experience with Kubernetes and SLURM for GPU-cluster job orchestration.
- Proficiency in Python and ML frameworks, with PyTorch strongly preferred.
- Experience with cloud GPU infrastructure, ideally CoreWeave or a similar GPU/HPC-focused cloud.
- Strong project-scoping, problem-solving, analytical, collaboration, and communication skills, particularly in research settings.
- Experience with geometric deep learning, neural operators, mesh/point-cloud/graph architectures, HPC simulation workflows, model serving, experiment tracking, observability, or customer deployment packaging is preferred.
Benefits
- Hybrid work model combining Shoreditch office time with work-from-home days.
- Equity options and a 10% employer pension contribution.
- Free office lunches, enhanced parental leave, YellowNest nursery scheme, and 25 days of annual leave plus public holidays.
- Private medical insurance, Wellhub subscription, eye tests, personal development support, and an Employee Assistance Programme.
- Bike2Work scheme, season ticket loan, and Octopus EV salary sacrifice.
Tech Stack
Categories
About PhysicsX
PhysicsX builds an AI-driven simulation software stack that enables high-fidelity, multi-physics modeling and optimization for engineering and manufacturing teams. The company sells software and delivery services to enterprises in aerospace and defense, automotive, semiconductors, materials, and energy. It is privately held and headquartered in London, with offices in London and New York.
