Firmus Technologies

AI Engineer - Inference

Firmus Technologies
Apply
7 hours ago
Sydney, AustraliaSenior

Responsibilities

  • Build, operate, and improve self-hosted AI inference services for internal applications, customer-facing products, and future Inference-as-a-Service offerings.
  • Define model-onboarding workflows covering compatibility validation, packaging, runtime selection, optimization, deployment, endpoint registration, testing, release, and lifecycle management.
  • Provision secure, scalable endpoints for generation, RAG, embeddings, reranking, batch processing, multimodal inference, tool calling, and agentic workflows.
  • Develop deployment templates, APIs, SDKs, configuration standards, and self-service workflows for managing model endpoints.
  • Optimize serving performance through quantization, compilation, batching, caching, routing, load balancing, memory optimization, and distributed parallelism.
  • Create benchmarked, versioned inference recipes covering models, runtimes, precision formats, GPU configurations, topology, scaling, scheduler profiles, and expected performance.
  • Design distributed inference configurations and collaborate with Kubernetes and scheduler teams on resource profiles, placement, quotas, autoscaling, and workload policies.
  • Build benchmarking, qualification, regression-testing, observability, and release workflows for models, runtimes, infrastructure, schedulers, and GPU platforms.
  • Measure and improve latency, throughput, concurrency, GPU and memory utilization, scaling efficiency, power efficiency, cost efficiency, and reliability.
  • Provide governed inference endpoints and performance information for agentic applications and contribute workload profiles to Model-to-Grid scheduling and capacity planning.
  • Collaborate with Product, UX, DevOps, Platform, Infrastructure, Security, and Global Operations teams.

Requirements

  • 5+ years of software engineering experience, including 3+ years in AI inference, model serving, ML systems, high-performance computing, distributed systems, or comparable performance-critical environments.
  • Demonstrated experience building, operating, or materially improving production model-serving platforms, inference APIs, GPU-backed services, AI developer platforms, or multi-tenant AI systems.
  • Hands-on experience with one or more modern inference frameworks, including TensorRT-LLM, TensorRT, SGLang, vLLM, Triton Inference Server, NVIDIA Dynamo, NVIDIA NIM, or Hugging Face Text Generation Inference.
  • Strong understanding of CUDA, cuDNN, NCCL, TensorRT, GPU profiling, distributed communication, and GPU performance analysis.
  • Practical understanding of LLM and generative-AI serving behavior, including token generation, batching, context length, concurrency, KV-cache management, scheduling, routing, and latency-throughput trade-offs.
  • Experience with quantization, compilation, calibration, mixed precision, kernel fusion, memory optimization, caching, speculative decoding, parallelism, and accuracy-performance validation.
  • Strong Python skills and working proficiency in C++ or Go for inference services, APIs, automation, benchmarking, profiling, runtime integrations, and performance-critical development.
  • Experience with distributed inference or training patterns, collective communication, fault handling, and multi-node scaling.
  • Familiarity with Kubernetes, containers, CI/CD, GitOps, service APIs, autoscaling, workload scheduling, observability, and production multi-tenant platform operations.
  • Understanding of GPU topology and high-performance infrastructure, including NVLink, NVSwitch, PCIe, NUMA, NIC affinity, RDMA, RoCEv2, network fabrics, and storage throughput.
  • Experience with inference benchmarking, profiling, reproducibility, load testing, regression testing, and performance analysis.
  • Understanding of security and governance for inference services, including identity, authentication, authorization, tenant isolation, quotas, rate limiting, secrets handling, audit logging, abuse prevention, and data protection.
  • Familiarity with RAG, embeddings, reranking, multimodal inference, agentic applications, model routing, and tool-calling workflows.

Benefits

  • Permanent full-time employment.
  • Location options in Singapore or Australia, including Launceston, Hobart, Sydney, or Melbourne.
  • Opportunity to work with founders and experts in AI infrastructure, energy systems, and next-generation compute.
  • Direct ownership and exposure to large-scale sustainable AI infrastructure and NVIDIA Cloud and Engineering partner technologies.
Firmus Technologies

About Firmus Technologies

51-200 employees
Contact me