Telnyx

Inference Infrastructure Architect (Remote)

Telnyx
Apply
9 hours ago
Remote, ChinaSenior / Staff+

Responsibilities

  • Architect and operate efficient GPU inference fleets while meeting latency and reliability SLOs and reducing cost per token.
  • Build serverless serving pools using vLLM and SGLang with batching, caching, quantization, and expert parallelism.
  • Develop the Kubernetes fleet layer for routing, prefill/decode disaggregation, KV-cache tiering, multi-LoRA serving, scheduling, isolation, and bare-metal lifecycle management.
  • Build dedicated per-tenant inference deployments with metering, private networking, adapter versioning, canary rollouts, and rollback capabilities.
  • Implement model distribution, warm pools, and inference-metric-driven autoscaling.
  • Operate observability and capacity-planning systems using engine metrics, DCGM, Prometheus, OpenTelemetry, and roofline analysis.
  • Contribute upstream, write runbooks and design documents, and help establish the founding China team.

Requirements

  • Owned production LLM serving under meaningful traffic and latency constraints and can explain bottlenecks, interventions, and measured outcomes.
  • Hands-on experience operating Kubernetes on GPU fleets, including GPU Operator, device plugins, node pools, topology-aware placement, gang scheduling, and GitOps rollouts.
  • Deep operational command of vLLM or SGLang, including deployment, tuning, upgrades, parallelism, quantization, batching, KV-cache configuration, prefix caching, and disaggregation.
  • Strong system-level performance engineering skills using engine and GPU metrics, roofline reasoning, and quantitative deployment sizing.
  • Proficiency with Python and Go plus substantial Linux, networking, and storage experience.
  • Ability to collaborate in English and Chinese and communicate through clear runbooks and design documents.
  • Experience with large-scale inference platforms or heavy production use of relevant open-source projects is especially valued.
  • Experience with LoRA and evaluation infrastructure, real-time voice latency, or multi-region and data-residency deployments is beneficial.

Benefits

  • Remote work from mainland China with no relocation required and a global, async-friendly team.
  • Opportunity to work on and shape a greenfield B300 GPU inference platform and help build the founding China team.
  • Open-source-first culture with supported conference travel and upstream contribution as part of the role.
  • Visa sponsorship is available if the employee later chooses to move to one of Telnyx’s hiring-entity locations.

Tech Stack

AmbassadorGoGrafanaKubernetesLinuxOpenStackPrometheusPython
Telnyx

About Telnyx

201-500 employees

Telnyx builds a communications platform and AI infrastructure for developers and enterprises, offering SIP trunking, voice and messaging APIs, WhatsApp Business, eSIM/IoT connectivity, and low-latency inference on a private global IP network. The privately held company, founded in 2009 and headquartered in Austin, operates worldwide and holds a Saudi VVSP license with regional anchorsites for data locality. Revenue comes from usage-based APIs and enterprise connectivity services.

Contact me