19 days ago
Atlanta, GA, USASenior
H1B sponsor

Responsibilities

  • Architect, implement, and manage enterprise-wide AI training and inference platform infrastructure.
  • Drive innovation, operational excellence, scalability, inference optimization, and model-serving performance.
  • Lead technical strategy and roadmaps for AI platform development, compute resource management, model lifecycle optimization, and real-time serving.
  • Guide cross-functional AI/ML engineering teams and mentor MLOps and platform engineers.
  • Align AI platform capabilities with model requirements, business SLAs, resource utilization goals, and cost-effective scaling.
  • Present training efficiency metrics, inference latency improvements, infrastructure costs, model benchmarks, and scalability roadmaps to technical and executive audiences.

Requirements

  • Extensive experience designing and managing AI/ML training and inference platforms using AWS, Azure, or GCP.
  • Deep expertise with ML model-serving frameworks, inference engines, distributed training, model optimization, and real-time model-serving architectures.
  • Experience managing GPU clusters, Kubernetes ML workloads, Docker containers, MLOps pipelines, model versioning, A/B testing, and continuous integration for ML models.
  • Advanced Python programming skills and experience with TensorFlow, PyTorch, Hugging Face, and related ML libraries.
  • Experience with high-performance computing, inference optimization, AI model monitoring, performance tracking, and observability.
  • Proven leadership and mentoring experience with cross-functional AI/ML, MLOps, and platform engineering teams.
  • Strong collaboration, strategic thinking, problem-solving, written communication, and verbal communication skills.
  • Advanced degree or extensive relevant experience in Computer Science, Machine Learning, Data Engineering, or a related AI/ML field is preferred.
  • Experience with scikit-learn, Transformers, vector databases, model registries, feature stores such as Feast or Tecton, Spark, Ray, CUDA, Prometheus, Grafana, AWS SageMaker, Azure ML, Google AI Platform, Helm, and automated ML retraining workflows is preferred.
  • Experience with enterprise AI/ML governance, compliance, and responsible AI practices is preferred.

Tech Stack

Apache SparkAWSAzureDockerGoogle Cloud PlatformGrafanaHelmKubernetesMLflowPrometheusPythonPyTorchscikit-learnTensorFlow
Intercontinental Exchange

About Intercontinental Exchange

10,000+ employees

ICE (NYSE: ICE) connects people to data, technology and expertise that create opportunity and inspire innovation. For terms of use, visit www.ice.com/privacy-security-center/terms-of-use

Contact me