Ernst and Young

Senior Associate/Manager - Applied AI Engineer, Technology Consulting

Ernst and Young
Apply
10 days ago
Singapore, SingaporeStaff+

Responsibilities

  • Own the architecture of production AI systems, including inference stacks, fine-tuning pipelines, retrieval and evaluation infrastructure, and monitoring.
  • Build production applications using frontier models such as Claude and GPT, including tool use, structured outputs, context and cost management, evaluations, and guardrails.
  • Deploy and operate open-source models such as Llama, Qwen, Mistral, and DeepSeek using quantization, model-serving frameworks, and multi-GPU inference.
  • Design and manage cloud infrastructure for GPU orchestration, autoscaling, cost controls, networking, IAM, and observability.
  • Fine-tune, distill, and evaluate models against real task metrics and determine which research advances are ready for production.
  • Contribute as a senior individual contributor or lead and mentor a small team depending on project needs.
  • Partner with product and leadership to translate ambiguous business problems into production systems and communicate technical tradeoffs to non-technical stakeholders.

Requirements

  • 6+ years of software or infrastructure engineering experience with deep production experience on AWS, GCP, or Azure.
  • Strong command of GPU infrastructure, including instance selection, drivers and CUDA, containerization, Kubernetes or an equivalent orchestrator, and inference autoscaling.
  • Experience with Terraform, Pulumi, or CDK; CI/CD; Prometheus, Grafana, or OpenTelemetry; monitoring; and cost management.
  • Fluency in Python and comfort with at least one systems-adjacent language such as Go, Rust, or C++.
  • Deep understanding of transformer internals, including attention variants, positional encodings, tokenization, KV caches, and sampling.
  • Production experience with frontier models, including tool use or function calling, structured outputs, retrieval, context and cost management, evaluations, and guardrails.
  • Hands-on experience fine-tuning open-source LLMs using full fine-tuning, LoRA, QLoRA, DPO, ORPO, or equivalent preference optimization methods.
  • Practical familiarity with PyTorch, Hugging Face, DeepSpeed or FSDP, vLLM or an equivalent serving framework, and evaluation frameworks.
  • Ability to read current research papers, assess their practical novelty, and determine whether they should be integrated into production.
  • Track record of deploying frontier or open-weight models into production systems with evaluation, guardrails, and monitoring.

Benefits

  • Flexible working environment and participation in globally connected, diverse, and inclusive teams.
  • Opportunity to work across EY’s consulting, technology, AI, and multidisciplinary client services.
  • Opportunity to combine hands-on individual-contributor work with architecture ownership, team leadership, and mentoring.
Ernst and Young

About Ernst and Young

10,000+ employees

Ernst & Young (EY) provides audit/assurance, tax, consulting, strategy and transactions services to enterprises, financial institutions, and public‑sector clients. Structured as a global network of partner‑owned member firms, it sells professional services on a fee basis, including a dedicated Financial Services Organization for banking, insurance, and capital markets. Headquartered in London, EY was formed in 1989 from the merger of Ernst & Whinney and Arthur Young, and operates in 150+ countries.

Contact me