
AI Systems Engineer - DevOps& Observability - Senior
Ernst and Young11 days ago
Base Salary
$107k - $177k/yr
Responsibilities
- Build and operate CI/CD/CV pipelines for AI services, agents, and runtime components across multiple environments.
- Operate production-scale model-serving and inference systems on GPU infrastructure.
- Own AI asset governance and discovery, including registries, model metadata, lineage, license management, and CVE/SBOM scanning.
- Manage quotas, rate limits, utilization, and per-tenant cost attribution for AI workloads.
- Own observability for metrics, logs, traces, dashboards, LLM debugging, evaluation, alerting, and SLA notifications.
- Operate the OpenTelemetry collection layer, GPU telemetry, exporters, queues, batching, and dynamic filtering.
- Automate GitOps delivery, continuous verification, progressive rollout, and rollback with policy, quality, integrity, and cost gates.
- Ensure AI workload consumption and telemetry are identity-stamped and attributable per tenant.
- Define ownership boundaries and consumption contracts with platform, trust, and data teams.
Requirements
- At least 8 years of experience in DevOps, MLOps, platform, or observability engineering.
- Hands-on production ownership of AI or high-throughput services.
- Strong DevOps experience with CI/CD/CV pipelines, GitOps, automated release, and rollback tooling.
- Hands-on experience operating Ray Serve, vLLM, Triton, or NIM on GPU infrastructure.
- Strong experience with Prometheus, Grafana, Loki, Tempo or Jaeger, and OpenTelemetry.
- Experience with API gateways and request routing, including streaming responses.
- Experience with OpenCost, Kubecost, or equivalent FinOps tooling and quota/rate-limit enforcement.
- Familiarity with model and artifact registries, supply-chain scanning, SBOMs, and license or lineage tracking.
- Experience operating AI or service infrastructure under compliance, security, or regulatory constraints.
- Ability to communicate runtime, cost, and observability tradeoffs to engineers, architects, and leadership.
- A bachelor’s or master’s degree in Computer Science or a related technical field is preferred.
- Experience with LangSmith or Langfuse, prompt and response quality measurement, sandboxed execution, GPU telemetry optimization, OpenLineage, AI license management, multi-tenant cost attribution, per-tenant SLA alerting, and regulated delivery environments is preferred.
Benefits
- Base salary is $106,900 to $176,500 across US geographic locations, with a stated range of $128,400 to $200,600 for New York City Metro Area, Washington State, and California excluding Sacramento.
- Medical and dental coverage, pension, 401(k) plans, and paid time off are offered.
- The role follows EY’s team-led, leader-enabled hybrid model, with most external client-serving employees expected to work in person 40–60% of the time over an engagement, project, or year.
- Flexible vacation, paid holidays, winter and summer breaks, personal and family care leave, and other leaves of absence are available.
- The position accepts applications on an ongoing basis and is located anywhere in the country.