
Staff AI Platform Engineer
Code Metal19 hours ago
Remote, United States +3 moreStaff+
Base Salary
$205k - $250k/yr
Responsibilities
- Set the technical direction and architecture for the AI platform, lead the four-person team, own design documents and RFCs, guide build-versus-buy decisions, and mentor engineers.
- Deploy, benchmark, and tune production inference for open-weight models on vLLM, SGLang, and TensorRT-LLM.
- Own the model gateway for self-hosted and commercial models, including authentication, routing, failover, quotas, and cost attribution.
- Build reusable agent harnesses and orchestration primitives for reliable and verifiable workflows.
- Develop context-engineering services for agent memory, retrieval, and data discovery.
- Instrument services with OpenTelemetry traces and metrics and build experiment-tracking and artifact infrastructure.
- Design the platform for multi-tenancy, versioned APIs, security, customer deployments, and air-gapped environments.
- Partner with Applied AI Research, product teams, and DevOps to align the platform with organizational needs.
Requirements
- Production-grade Python and strong platform engineering fundamentals, including API and service design, distributed systems, containers, Kubernetes, CI/CD, and testing.
- Experience shipping production agentic systems and improving their reliability.
- Experience with retrieval-augmented generation, embeddings, vector or hybrid search, and agent memory, ideally involving code or large technical corpora.
- Experience instrumenting services with tracing and metrics and operating AI services against SLOs.
- Data science and AI research fundamentals covering transformers, LLM inference, experiment design, benchmarking, and model evaluation.
- Working familiarity with PyTorch and Hugging Face, plus experience fine-tuning, evaluating, or serving language models.
- Staff-level technical leadership across multiple systems or teams, including architecture ownership, design documents and RFCs, roadmap development, and mentoring.
- Typically 8+ years of software engineering experience, including 4+ years building and operating ML or LLM systems in production; equivalent depth of experience is valued.
- Preferred experience with LLM gateways or proxies such as SMG or Bifrost, or equivalent API gateway development.
- Preferred experience optimizing inference through speculative decoding, prefix caching, tensor/pipeline/expert parallelism, disaggregated prefill and decode, and GPU profiling.
- Preferred experience building LLM and agent evaluation harnesses and experiment-tracking or artifact systems such as MLflow or Weights & Biases.
- Preferred experience taking internal platforms to external products, deploying AI systems in on-premises, air-gapped, classified, regulated, defense, or aerospace environments.
Benefits
- Health care plan with 100% premium coverage, including medical, dental, and vision.
- 401(k) with 5% matching.
- Uncapped vacation, sick leave, and public holidays.
- Flexible hybrid or remote work arrangement.
- Relocation assistance for qualifying employees.
- US citizenship may be required for certain security-clearance project assignments.
Tech Stack
Categories
About Code Metal
Code Metal builds an AI-assisted code transpilation and verification platform that translates algorithms into optimized implementations for heterogeneous hardware (GPUs, FPGAs, ASICs, and edge SoCs). It serves defense, automotive, aerospace, and semiconductor teams deploying DSP, RF, communications, and embedded signal-processing workloads to constrained devices. Privately held and founded in 2023, the company is headquartered in Boston, MA.