3 months ago
Responsibilities
- Design and operate distributed training infrastructure for neural operator architectures on NVIDIA DGX B200 systems.
- Optimize training throughput, fault tolerance, and cost efficiency through checkpointing, gradient accumulation, and multi-node synchronization.
- Build and maintain experiment tracking and observability systems for training runs, hyperparameter sweeps, and model performance.
- Resolve data-loading bottlenecks for large-scale mesh datasets and optimize cloud-storage I/O through prefetching, caching, and format optimization.
- Build serving infrastructure for pre-trained large physics models supporting zero-shot inference and uncertainty quantification.
- Design model-packaging pipelines for reliable customer deployment with fine-tuning capabilities.
- Ensure model checkpoints are reproducibly deployable with consistent behavior.
- Improve research developer experience through faster iteration, reliable delivery pipelines, and debugging tools.
- Collaborate with the broader Infrastructure team on shared infrastructure patterns and standards.
Requirements
- 5+ years of experience building and operating ML infrastructure at scale.
- Deep expertise in distributed training, including debugging NCCL hangs and optimizing collective communication with FSDP, DDP, and pipeline parallelism.
- Strong systems fundamentals across Linux, networking, NVLink, InfiniBand, storage I/O, profiling, and performance optimization.
- Production experience with Kubernetes and SLURM for orchestration on GPU clusters.
- Proficiency in Python and machine learning frameworks, with PyTorch strongly preferred.
- Experience with cloud GPU infrastructure; CoreWeave or a similar GPU/HPC-focused cloud is preferred.
- Ability to scope and deliver projects while prioritizing work effectively.
- Strong problem-solving, collaboration, and communication skills, particularly in research environments.
- Experience with geometric deep learning or neural operators working on meshes, point clouds, or graphs is preferred.
- Background in HPC for simulation engineering and familiarity with CFD/FEA data workflows is preferred.
- Experience building model-serving infrastructure with latency and throughput requirements is preferred.
- Familiarity with Weights & Biases, MLflow, Prometheus, and Grafana is preferred.
- Experience packaging models for customer deployment, including containers, model registries, and versioning, is preferred.
Benefits
- Hybrid work model combining Shoreditch office time with work-from-home days.
- Equity options.
- 10% employer pension contribution.
- Free office lunches.
- Enhanced parental leave, including 3 months full-pay paternity leave and 6 months full-pay maternity leave.
- YellowNest nursery scheme.
- 25 days of annual leave plus public holidays.
- Private medical insurance with 100% employee cover.
- Wellhub subscription.
- Eye tests.
- Personal development support.
- Employee Assistance Programme.
- Bike2Work scheme and season ticket loan.
- Octopus EV salary sacrifice scheme.
Tech Stack
Categories
About PhysicsX
PhysicsX is a physical AI company on a mission to accelerate innovation and overhaul what engineering and manufacturing look like today. We are building a new software stack to deliver deep AI enablement across the entire engineering lifecycle. PhysicsX partners with leading organizations in aerospace & defense, automotive, semiconductors, materials, and energy, supporting them on some of their most critical and complex challenges. PhysicsX is headquartered in the United Kingdom, with offices in London and New York. We are currently recruiting for multiple positions, however, please only apply for the role that best aligns with your skillset and career goals.
