Meta

Software Engineer, GenAI Frameworks - MTIA

Meta
Apply
1 day ago
Menlo Park, CA, USASenior
H1B sponsor

Base Salary

$184k - $257k/yr

Responsibilities

  • Own a key GenAI inference framework area covering serving runtime, distributed inference, graph-mode execution and compilation, or core PyTorch integration.
  • Lead ambiguous multi-quarter technical programs across teams, including design, implementation, testing, CI, rollout, and production hardening.
  • Design, implement, and ship generative AI inference framework features and MTIA integration layers from prototype through production.
  • Profile and optimize production inference across frameworks, runtimes, compilers, and kernels to improve latency, throughput, and resource efficiency.
  • Build distributed inference capabilities including hierarchical KV caching, cross-host transfers, disaggregated prefill and decode, and parallelism strategies.
  • Optimize graph capture and replay, host-side latency, memory allocation and placement, and execution under real traffic.
  • Enable frontier models on MTIA and validate accuracy against GPU baselines while resolving correctness regressions.
  • Own accuracy, stability, latency, throughput, and regression-test quality for the assigned area.
  • Partner with kernel, compiler, and silicon teams on hardware/software co-design.
  • Mentor engineers, lead design reviews, and communicate technical decisions across teams.

Requirements

  • Bachelor’s degree in Computer Science, Computer Engineering, a relevant technical field, or equivalent practical experience.
  • 8+ years of professional experience in systems software, ML infrastructure, performance engineering, or compiler/framework development.
  • 5+ years of hands-on experience with generative AI inference or training frameworks, or production model-serving systems.
  • Proficiency in Python, C++, or Rust, including low-level systems code and performance-critical paths.
  • Experience optimizing GenAI inference, including prefill and decode, KV-cache management, batching, scheduling, latency, and throughput.
  • Experience with deep learning framework internals such as operator registration and dispatch, eager execution, autograd, graph capture, and compilation.
  • Experience with runtime optimization involving graph-mode execution, host-side latency, memory allocation, and placement.
  • Experience leading technical design and end-to-end delivery of frameworks or infrastructure projects across team boundaries.
  • Experience working across framework, compiler, runtime, and hardware layers and debugging issues across component boundaries.
  • Experience diagnosing numerical accuracy issues and maintaining accuracy and stability standards.
  • Preferred experience with long-context inference, context and sequence parallelism, ring or streaming attention, and hierarchical or offloaded KV caches.
  • Preferred experience with GPU, TPU, or custom ASIC platforms and bringing up new hardware backends in major ML frameworks.
  • Preferred experience with low-precision numerics and quantization, including FP8, block-scaled formats, INT8, INT4, calibration, and accuracy analysis.
  • Preferred experience with distributed inference or training at scale, tensor, pipeline, and expert parallelism, collective communication, RDMA, and multi-host serving.
  • Familiarity with the PyTorch compilation stack, including TorchDynamo, FX IR, Torch Inductor, dynamic shapes, Triton, AOT I, and custom operators.
  • Preferred experience with LLM serving features such as continuous batching, paged or chunked attention, speculative decoding, prefix caching, disaggregated prefill and decode, and mixture-of-experts routing.
  • Experience using profiling and tracing to resolve production performance bottlenecks is preferred.
  • Preferred ability to use AI tools to improve workflows and familiarity with responsible AI practices, prompt/context engineering, and agent orchestration.

Benefits

  • Annual base salary of $183,997/year to $257,000/year, plus bonus, equity, and benefits.
Meta

About Meta

10,000+ employees

Meta builds social platforms and communication apps—including Facebook, Instagram, WhatsApp, and Messenger—and develops AR/VR hardware and software such as Quest to power immersive computing. It monetizes primarily through advertising tools for businesses, with additional revenue from devices and services, and operates a massive global infrastructure. Founded in 2004 and headquartered in Menlo Park, California, Meta Platforms, Inc. is a public company traded on Nasdaq under the ticker META.

Contact me