26 days ago
Remote, NetherlandsSenior
Responsibilities
- Lead the design and implementation of production-grade ML and Generative AI solutions on AWS with awareness of multi-cloud environments.
- Architect and optimize cloud systems for cost efficiency, reliability, resilience, security, performance, and reduced operational toil.
- Deliver cloud optimization sessions, performance and efficiency workshops, security posture reviews, reliability reviews, and architecture assessments.
- Support customer Expert Inquiry requests and provide deep cloud engineering resolutions.
- Deploy and operate ML and GenAI workloads, including training, inference, GPU utilization, scaling, monitoring, logging, and FinOps.
- Convert customer solutions into reusable playbooks, Terraform modules, CloudFlow templates, cloud diagrams, Composer Recipes, and documentation.
- Provide structured feedback to Product and Engineering teams and contribute code, feature requests, internal tooling, and DCI features.
- Build agent skills, scripts, and automation that scale cloud and AI/ML expertise across the team.
- Work as the embedded technical partner within account teams alongside Customer Success Managers and Account Managers.
- Use DoiT Cloud Intelligence products to create analytics and allocation dashboards, identify opportunities, build Composer recipes, and implement CloudFlow automations.
- Enable customers to integrate DCI with observability, CI/CD, and governance processes.
- Share knowledge through documentation, demos, office hours, training, design reviews, and internal forums.
Requirements
- 4+ years of experience architecting, deploying, and managing production cloud-based AI/ML solutions.
- Proven experience designing and operating large distributed systems on AWS.
- Advanced proficiency with AWS services relevant to AI/ML and Generative AI.
- Hands-on experience with Amazon Bedrock, Amazon SageMaker, SageMaker JumpStart, LLMs, multimodal AI, prompt engineering, and model evaluation.
- Knowledge of agentic AI patterns, Amazon Q Business, and Amazon Q Developer or similar tools.
- Experience with SageMaker Pipelines, Model Monitor, Data Wrangler, and SageMaker Clarify.
- Proficiency integrating TensorFlow and PyTorch with SageMaker for model development, fine-tuning, and deployment.
- Experience with distributed training, multi-GPU or multi-node systems, and inference performance optimization.
- Strong AWS data-engineering experience with Amazon S3, AWS Glue, Lake Formation, and Redshift.
- Experience building workflows with AWS Lambda, Step Functions, API Gateway, Amazon EKS, and AWS Fargate.
- Hands-on AI/ML CI/CD experience using AWS CodePipeline, CodeBuild, SageMaker Pipelines, or similar tools.
- Experience monitoring AI systems with Amazon CloudWatch and SageMaker Model Monitor.
- Understanding of AI governance, security, compliance, IAM, KMS, data privacy, AI ethics, and bias mitigation.
- Working knowledge of Google Cloud AI tools such as Vertex AI, Cloud AutoML, and BigQuery ML.
- Ability to mentor peers, run enablement sessions, collaborate across Sales, Customer Success, and Product, and communicate with technical and business audiences.
- A BA/BS degree in Computer Science, Mathematics, or a related technical field, or equivalent practical experience, is listed as a bonus qualification.
- Additional data or AI certifications, RLHF, advanced fine-tuning, hybrid AI architectures, Hugging Face, consulting or SaaS AI experience, and JIRA experience are bonus qualifications.
Benefits
- Remote-first work with flexible working options across the listed employee locations; contractors may be based in Eastern Europe or Portugal.
- Unlimited vacation.
- Health insurance.
- Parental leave.
- Employee stock option plan.
- Home office allowance.
- Professional development stipend.
- Peer recognition program.
Tech Stack
Categories
Forward Deployed
