AI Engineer - Inference
Firmus Technologies7 hours ago
Sydney, AustraliaSenior
Responsibilities
- Build, operate, and improve self-hosted AI inference services for internal applications, customer-facing products, and future Inference-as-a-Service offerings.
- Define model-onboarding workflows covering compatibility validation, packaging, runtime selection, optimization, deployment, endpoint registration, testing, release, and lifecycle management.
- Provision secure, scalable endpoints for generation, RAG, embeddings, reranking, batch processing, multimodal inference, tool calling, and agentic workflows.
- Develop deployment templates, APIs, SDKs, configuration standards, and self-service workflows for managing model endpoints.
- Optimize serving performance through quantization, compilation, batching, caching, routing, load balancing, memory optimization, and distributed parallelism.
- Create benchmarked, versioned inference recipes covering models, runtimes, precision formats, GPU configurations, topology, scaling, scheduler profiles, and expected performance.
- Design distributed inference configurations and collaborate with Kubernetes and scheduler teams on resource profiles, placement, quotas, autoscaling, and workload policies.
- Build benchmarking, qualification, regression-testing, observability, and release workflows for models, runtimes, infrastructure, schedulers, and GPU platforms.
- Measure and improve latency, throughput, concurrency, GPU and memory utilization, scaling efficiency, power efficiency, cost efficiency, and reliability.
- Provide governed inference endpoints and performance information for agentic applications and contribute workload profiles to Model-to-Grid scheduling and capacity planning.
- Collaborate with Product, UX, DevOps, Platform, Infrastructure, Security, and Global Operations teams.
Requirements
- 5+ years of software engineering experience, including 3+ years in AI inference, model serving, ML systems, high-performance computing, distributed systems, or comparable performance-critical environments.
- Demonstrated experience building, operating, or materially improving production model-serving platforms, inference APIs, GPU-backed services, AI developer platforms, or multi-tenant AI systems.
- Hands-on experience with one or more modern inference frameworks, including TensorRT-LLM, TensorRT, SGLang, vLLM, Triton Inference Server, NVIDIA Dynamo, NVIDIA NIM, or Hugging Face Text Generation Inference.
- Strong understanding of CUDA, cuDNN, NCCL, TensorRT, GPU profiling, distributed communication, and GPU performance analysis.
- Practical understanding of LLM and generative-AI serving behavior, including token generation, batching, context length, concurrency, KV-cache management, scheduling, routing, and latency-throughput trade-offs.
- Experience with quantization, compilation, calibration, mixed precision, kernel fusion, memory optimization, caching, speculative decoding, parallelism, and accuracy-performance validation.
- Strong Python skills and working proficiency in C++ or Go for inference services, APIs, automation, benchmarking, profiling, runtime integrations, and performance-critical development.
- Experience with distributed inference or training patterns, collective communication, fault handling, and multi-node scaling.
- Familiarity with Kubernetes, containers, CI/CD, GitOps, service APIs, autoscaling, workload scheduling, observability, and production multi-tenant platform operations.
- Understanding of GPU topology and high-performance infrastructure, including NVLink, NVSwitch, PCIe, NUMA, NIC affinity, RDMA, RoCEv2, network fabrics, and storage throughput.
- Experience with inference benchmarking, profiling, reproducibility, load testing, regression testing, and performance analysis.
- Understanding of security and governance for inference services, including identity, authentication, authorization, tenant isolation, quotas, rate limiting, secrets handling, audit logging, abuse prevention, and data protection.
- Familiarity with RAG, embeddings, reranking, multimodal inference, agentic applications, model routing, and tool-calling workflows.
Benefits
- Permanent full-time employment.
- Location options in Singapore or Australia, including Launceston, Hobart, Sydney, or Melbourne.
- Opportunity to work with founders and experts in AI infrastructure, energy systems, and next-generation compute.
- Direct ownership and exposure to large-scale sustainable AI infrastructure and NVIDIA Cloud and Engineering partner technologies.