
Senior Associate/Manager - Applied AI Engineer, Technology Consulting
Ernst and Young10 days ago
Singapore, SingaporeStaff+
Responsibilities
- Own the architecture of production AI systems, including inference stacks, fine-tuning pipelines, retrieval and evaluation infrastructure, and monitoring.
- Build production applications using frontier models such as Claude and GPT, including tool use, structured outputs, context and cost management, evaluations, and guardrails.
- Deploy and operate open-source models such as Llama, Qwen, Mistral, and DeepSeek using quantization, model-serving frameworks, and multi-GPU inference.
- Design and manage cloud infrastructure for GPU orchestration, autoscaling, cost controls, networking, IAM, and observability.
- Fine-tune, distill, and evaluate models against real task metrics and determine which research advances are ready for production.
- Contribute as a senior individual contributor or lead and mentor a small team depending on project needs.
- Partner with product and leadership to translate ambiguous business problems into production systems and communicate technical tradeoffs to non-technical stakeholders.
Requirements
- 6+ years of software or infrastructure engineering experience with deep production experience on AWS, GCP, or Azure.
- Strong command of GPU infrastructure, including instance selection, drivers and CUDA, containerization, Kubernetes or an equivalent orchestrator, and inference autoscaling.
- Experience with Terraform, Pulumi, or CDK; CI/CD; Prometheus, Grafana, or OpenTelemetry; monitoring; and cost management.
- Fluency in Python and comfort with at least one systems-adjacent language such as Go, Rust, or C++.
- Deep understanding of transformer internals, including attention variants, positional encodings, tokenization, KV caches, and sampling.
- Production experience with frontier models, including tool use or function calling, structured outputs, retrieval, context and cost management, evaluations, and guardrails.
- Hands-on experience fine-tuning open-source LLMs using full fine-tuning, LoRA, QLoRA, DPO, ORPO, or equivalent preference optimization methods.
- Practical familiarity with PyTorch, Hugging Face, DeepSpeed or FSDP, vLLM or an equivalent serving framework, and evaluation frameworks.
- Ability to read current research papers, assess their practical novelty, and determine whether they should be integrated into production.
- Track record of deploying frontier or open-weight models into production systems with evaluation, guardrails, and monitoring.
Benefits
- Flexible working environment and participation in globally connected, diverse, and inclusive teams.
- Opportunity to work across EY’s consulting, technology, AI, and multidisciplinary client services.
- Opportunity to combine hands-on individual-contributor work with architecture ownership, team leadership, and mentoring.
Tech Stack
Categories
About Ernst and Young
Ernst & Young (EY) provides audit/assurance, tax, consulting, strategy and transactions services to enterprises, financial institutions, and public‑sector clients. Structured as a global network of partner‑owned member firms, it sells professional services on a fee basis, including a dedicated Financial Services Organization for banking, insurance, and capital markets. Headquartered in London, EY was formed in 1989 from the merger of Ernst & Whinney and Arthur Young, and operates in 150+ countries.