Accellor

AI Principal Engineer

Accellor
Apply
1 month ago
Mountain View, CA, USA or San Francisco, CA, USAStaff+

Responsibilities

  • Design and evolve large-scale AI systems spanning inference runtime, model serving, GPU scheduling, distributed execution, observability, release gates, and production rollout.
  • Architect and optimize high-throughput, low-latency inference and model-serving systems, including batching, caching, routing, streaming, tensor parallelism, pipeline parallelism, and model sharding.
  • Analyze GPU kernels, memory movement, collective communication, scheduling, networking, and distributed execution to improve performance and cost efficiency.
  • Design context-engineering frameworks for prompts, retrieval, conversation memory, agent state, multimodal context, grounding, permissions, and context compression.
  • Build cost-optimization frameworks using token budgeting, prompt compression, semantic and response caching, model routing, batching, asynchronous execution, and fallback strategies.
  • Support distributed training, post-training, reinforcement learning, checkpointing, evaluation infrastructure, fault tolerance, and experiment execution.
  • Define release validation and evaluation gates covering correctness, latency, throughput, cost, context quality, retrieval quality, safety, reliability, and model output quality.
  • Design telemetry, tracing, dashboards, alerts, profiling, runbooks, SLOs, and production learning loops for AI infrastructure.
  • Support reliable and observable agentic, tool-use, memory, function-calling, multimodal, and long-running workflow platforms.
  • Provide technical leadership through architecture documents, RFCs, technical reviews, mentoring, and cross-functional architecture decisions.

Requirements

  • 10–12 years of experience in software engineering, systems architecture, ML infrastructure, distributed systems, platform engineering, inference systems, cloud infrastructure, or large-scale backend engineering.
  • Strong hands-on Python experience plus at least one of C++, Go, Rust, Java, or TypeScript.
  • Deep understanding of distributed systems, production infrastructure, reliability engineering, scalability, observability, and fault-tolerant architecture.
  • Experience designing or operating systems involving APIs, microservices, distributed compute, orchestration, job scheduling, caching, high availability, and production monitoring.
  • Strong knowledge of AI/ML systems, model serving, inference workflows, context engineering, retrieval systems, evaluation pipelines, and production model deployment.
  • Practical experience with GPU systems, CUDA/Triton-style programming, distributed inference, GPU profiling, memory optimization, and NCCL or RCCL.
  • Experience with ML frameworks and serving stacks such as PyTorch, JAX, TensorFlow, Triton, vLLM-style serving, Apache Ray, Kubernetes-based serving, or internal model-serving systems.
  • Ability to debug issues across model behavior, runtime systems, distributed infrastructure, networking, GPU execution, context quality, retrieval quality, evaluation harnesses, and production services.
  • Preferred experience includes LLM or multimodal inference, agent infrastructure, frontier-model serving, tensor and pipeline parallelism, KV-cache optimization, speculative decoding, long-context serving, context platforms, model routing, semantic caching, and LLM cost dashboards.
  • Preferred experience includes GPU profiling with Nsight Systems, Nsight Compute, rocprof, perf, Prometheus, Grafana, or OpenTelemetry; distributed training; RL infrastructure; ML compiler optimization; release gates; canary systems; regression detection; evals; grounding evaluation; safety testing; and model behavior monitoring.
  • Strong communication skills and ability to write architecture documents, assess trade-offs, review implementation quality, and align teams around technical decisions.

Benefits

  • Hybrid workplace in San Francisco, United States.
  • Delivery-department role with cross-functional collaboration across Research, Inference, Runtime, Infrastructure, Product, Safety, Security, Technical Success, and Deployment teams.
Accellor

About Accellor

201-500 employees
Contact me