Dialpad

Sr. Software Engineer, ML Inference Platform

Dialpad
Apply
3 hours ago
Buenos Aires, ArgentinaSenior

Responsibilities

  • Design, build, and improve shared capabilities spanning model training, evaluation, artifact management, release, inference, and operational feedback.
  • Build and operate reliable GPU training infrastructure, including scheduling, workload isolation, capacity management, storage, networking, observability, and accelerator utilization.
  • Develop low-latency, high-throughput, highly available production-serving pathways for inference workloads.
  • Integrate model-training frameworks and inference runtimes for automation, observability, security, and operational control.
  • Optimize GPU workloads across compute, memory, storage, networking, batching, concurrency, and scheduling.
  • Partner with ASR and NLP scientists on reproducibility, evaluation, scalability, hardware constraints, latency, reliability, cost, and release safety.
  • Improve model and artifact versioning, validation, promotion, deployment, rollback, benchmarking, evaluation, telemetry, logging, tracing, dashboards, alerting, and diagnostics.
  • Lead technical projects through production, contribute to architecture, mentor engineers, and build self-service workflows and standards for AI teams.

Requirements

  • Seven or more years of professional software engineering experience with ownership of backend, infrastructure, distributed, or ML platform systems in production.
  • Experience building or operating systems that support model training, model inference, or the lifecycle connecting them.
  • Proficiency in Python, Go, or another backend-oriented language, with experience delivering maintainable production software and well-designed interfaces.
  • Hands-on experience with Linux, containers, Kubernetes, cloud infrastructure, deployment automation, and production operations.
  • Experience operating GPU workloads and analyzing utilization, memory, storage, networking, scheduling, and workload performance.
  • Working knowledge of datasets, experiments, distributed execution, checkpoints, reproducibility, and model artifacts.
  • Understanding of dataset quality, evaluation design, experimental validity, error analysis, model-quality metrics, and production model behavior.
  • Ability to make trade-offs among model quality, latency, throughput, reliability, capacity, and cost, with strong operational judgment around observability and release safety.
  • Ability to lead ambiguous projects, communicate across disciplines, mentor engineers, and influence technical decisions.
  • Experience with ASR, speech processing, NLP, large language models, distributed model training, model-serving runtimes, PyTorch, JAX, Kubernetes-based GPU scheduling, GCP infrastructure, experiment tracking, model evaluation, artifact registries, production model monitoring, or internal platforms is particularly relevant.

Benefits

  • Competitive salary, comprehensive benefits, competitive benefits and perks, cutting-edge AI tools, and a robust training program.
  • Opportunities to work on agentic AI products and grow professionally.
  • Inclusive offices designed to support collaboration and connection.
  • Equal-opportunity workplace free from discrimination and harassment.
Dialpad

About Dialpad

1,001-5,000 employees

Dialpad is the leading AI-powered communications intelligence platform creating human-first, AI-enhanced solutions that will drive the next wave of how businesses communicate with and serve their customers. Enterprise customers such as Randstad, RE/MAX, Nasdaq, Express Scripts, T-Mobile, Johns Hopkins, Motorola Solutions, Tractor Supply, and Netflix use Dialpad and its AI capabilities to deliver amazing customer experiences. Supported by notable investors such as Andreessen Horowitz, GV, ICONIQ Capital, and OMERS, Dialpad is a dynamic force in AI technology with a rapidly expanding presence. Visit dialpad.com to learn more.

Contact me