2 months ago
Responsibilities
- Design, build, and improve systems connecting AI capability development to production inference.
- Build and optimize low-latency, high-throughput, high-availability model-serving and runtime systems.
- Operate and optimize containerized Kubernetes workloads in GCP, including efficient NVIDIA GPU, memory, storage, and networking utilization.
- Integrate model-serving frameworks such as vLLM, Triton, and TGI with internal deployment, observability, and release requirements.
- Enable shadow serving, canary rollouts, staged deployments, candidate-versus-incumbent comparisons, and fast rollback mechanisms.
- Build benchmarking and evaluation tools for latency, throughput, cost, saturation behavior, and reliability under realistic production traffic.
- Improve packaging, versioning, promotion, deployment, and rollback processes for model and capability artifacts.
- Strengthen runtime telemetry, structured logging, tracing, dashboards, and alerting for production model-serving systems.
- Improve compute efficiency, GPU utilization, autoscaling behavior, and cost-performance tradeoffs across the inference platform.
Requirements
- 6+ years of professional software engineering experience shipping backend services, infrastructure systems, or production platforms.
- Proficiency in Python, Go, or another backend-oriented programming language, with strong debugging and systems-thinking skills.
- Experience building, operating, or optimizing high-throughput services, distributed systems, data or ML infrastructure, or runtime platforms.
- Hands-on experience with containers, Kubernetes, Linux environments, deployment automation, and production operations.
- Ability to reason about compute, memory, networking, storage, batching, concurrency, latency, reliability, and service-level objectives.
- Strong judgment around reproducibility, observability, rollout safety, failure modes, and whole-system resilience.
- Ability to collaborate with model developers, product engineers, infrastructure teams, and technical leadership.
Benefits
- Competitive benefits and perks, including a robust training program and AI tools.
- Inclusive office environment designed to support collaboration and connection.
Categories
About Dialpad
Dialpad is the leading AI-powered communications intelligence platform creating human-first, AI-enhanced solutions that will drive the next wave of how businesses communicate with and serve their customers. Enterprise customers such as Randstad, RE/MAX, Nasdaq, Express Scripts, T-Mobile, Johns Hopkins, Motorola Solutions, Tractor Supply, and Netflix use Dialpad and its AI capabilities to deliver amazing customer experiences. Supported by notable investors such as Andreessen Horowitz, GV, ICONIQ Capital, and OMERS, Dialpad is a dynamic force in AI technology with a rapidly expanding presence. Visit dialpad.com to learn more.