1 day ago
Base Salary
$154k - $217k/yr
Responsibilities
- Design, implement, and optimize high-performance compute and communication kernels for MTIA accelerators from architectural analysis through production deployment.
- Profile and root-cause performance across instruction scheduling, memory hierarchy and DMA behavior, on-chip interconnect, and multi-device collectives.
- Build and extend kernel authoring frameworks, templates, and libraries that enable other engineers to achieve high performance.
- Deliver broad PyTorch operator coverage for recommendation, ranking, and generative AI workloads across eager and compiled execution paths.
- Partner with silicon architecture and design teams on hardware/software co-design, including pre-silicon roofline analysis and feature evaluation.
- Collaborate with compiler, runtime, framework, and product-facing teams to improve end-to-end production model performance.
- Investigate numerics and precision trade-offs and design software mitigations for hardware limitations.
- Set technical direction, write design documents, and mentor engineers on accelerator programming and performance methodology.
Requirements
- Bachelor’s degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience.
- At least 6 years of professional experience in high-performance computing, accelerator kernel development, compiler backends, or systems performance engineering.
- Proficiency in C++ and Python, including low-level systems programming, templates, generic programming, and performance-critical code.
- Experience writing and optimizing kernels for parallel architectures such as GPUs, TPUs, AI ASICs, or SIMD/vector CPUs.
- Working knowledge of computer architecture, including memory hierarchies, bandwidth, latency hiding, occupancy, scheduling, vectorization, and synchronization.
- Ability to build analytical or roofline performance models, profile against them, and explain residual performance gaps.
- Preferred: 8+ years of experience in accelerator software, HPC, or ML systems performance, or equivalent experience with an advanced degree.
- Preferred: experience with transformer and attention kernels, low-precision numerics and quantization, compiler/code-generation technologies, distributed execution, kernel libraries, pre-silicon development, hardware/software co-design, and ML framework internals.
- Preferred: experience mentoring engineers, setting technical direction, contributing to open source, applying responsible AI practices, and using AI tools to improve workflows.
Benefits
- Annual base salary of $154,003 to $217,000, plus bonus, equity, and benefits.
- The role covers pre-silicon simulation and emulation, first-silicon bring-up, and production deployment.
About Meta
Meta builds social platforms and communication apps—including Facebook, Instagram, WhatsApp, and Messenger—and develops AR/VR hardware and software such as Quest to power immersive computing. It monetizes primarily through advertising tools for businesses, with additional revenue from devices and services, and operates a massive global infrastructure. Founded in 2004 and headquartered in Menlo Park, California, Meta Platforms, Inc. is a public company traded on Nasdaq under the ticker META.
