26 days ago
Remote, EstoniaSenior
Responsibilities
- Lead the design and implementation of production-grade ML and Generative AI solutions on AWS, with awareness of multi-cloud environments.
- Advise customers on AI/ML workloads at scale, including discovery, deployment, optimization, training, inference, GPU utilization, scaling, monitoring, logging, and cost control.
- Design cloud architectures that improve cost efficiency, reliability, resilience, security, performance, observability, and operational efficiency.
- Deliver cloud optimization sessions, performance and efficiency workshops, security posture reviews, reliability reviews, and architecture deep dives.
- Respond to expert inquiries and support requests requiring deep cloud engineering expertise.
- Convert customer solutions into reusable playbooks, Terraform modules, CloudFlow templates, diagrams, recipes, and documentation.
- Provide product feedback and contribute code, features, agent skills, scripts, and internal tooling to DoiT Cloud Intelligence.
- Partner with Customer Success Managers, Account Managers, customer engineers, architects, and FinOps teams to execute technical deployments, integrations, automation, and platform adoption.
- Use DoiT Cloud Intelligence products to create dashboards, identify and resolve cost and reliability opportunities, build queries and insights, and implement automations and guardrails.
- Contribute to documentation, demos, office hours, training, design reviews, and enablement for FDE and Customer Success teams.
Requirements
- At least 4 years of experience architecting, deploying, and managing production cloud-based AI/ML solutions.
- Proven experience designing and operating large distributed systems on AWS.
- Advanced proficiency with AWS services relevant to AI/ML and Generative AI.
- Hands-on experience with Amazon Bedrock, Amazon SageMaker, LLMs, multimodal AI, prompt engineering, model evaluation, and agentic AI patterns.
- Knowledge of SageMaker Pipelines, Model Monitor, Data Wrangler, SageMaker Clarify, distributed training, inference optimization, and MLOps.
- Experience integrating TensorFlow and PyTorch with SageMaker for model development, fine-tuning, and deployment.
- Strong AWS data-engineering experience with Amazon S3, AWS Glue, Lake Formation, and Redshift.
- Experience building AI/ML workflows with AWS Lambda, Step Functions, API Gateway, Amazon EKS, and AWS Fargate.
- Experience with AI/ML CI/CD using AWS CodePipeline, AWS CodeBuild, SageMaker Pipelines, or similar tools.
- Knowledge of Amazon CloudWatch, IAM, KMS, AI governance, security, compliance, data privacy, and bias detection or mitigation.
- Working knowledge of Google Cloud AI tools such as Vertex AI, Cloud AutoML, and BigQuery ML.
- Ability to mentor peers, run enablement sessions, collaborate across Sales, Customer Success, and Product, and communicate with technical and business audiences.
- Preferred qualifications include a BA/BS in Computer Science, Mathematics, or a related technical field, or equivalent practical experience.
- Preferred experience includes RLHF, advanced fine-tuning, hybrid AI architectures, Hugging Face, consulting or SaaS delivery, JIRA, and Agile practices.
Benefits
- Remote-first work with flexible working options; full-time roles are available in the UK, Ireland, Estonia, Sweden, the Netherlands, and Israel, and contractor roles are available in Eastern Europe or Portugal.
- Unlimited vacation.
- Health insurance.
- Parental leave.
- Employee stock option plan.
- Home office allowance.
- Professional development stipend.
- Peer recognition program.
Tech Stack
Categories
Forward Deployed
