Staff AI Engineer
Blip Global1 month ago
Remote, BrazilStaff+
Responsibilities
- Lead the full lifecycle of language models and AI solutions, including design, training, validation, deployment, monitoring, and evolution.
- Evaluate and orchestrate transitions between commercial model APIs and internally distilled models deployed in a VPC.
- Conduct applied research and rigorous experiments involving architectures, quantization, and modeling strategies.
- Build and manage large-scale automated data cleaning and curation pipelines.
- Orchestrate and monitor scalable inference workloads in cloud environments.
- Partner with product and business teams to align AI initiatives with strategic company goals.
- Provide technical leadership, mentorship, and communication between AI research and business stakeholders.
Requirements
- Degree in Systems Engineering, Computer Science, Computer Engineering, Artificial Intelligence, or a related field.
- Practical mastery of knowledge distillation, Teacher-Student architectures, PEFT, LoRA, QLoRA, and adaptation of open models.
- Ability to compare and integrate proprietary model APIs with open models and deploy customized models.
- Experience with high-performance language-model serving, inference optimization, and quantization.
- Ability to extract, process, sanitize, and curate large datasets and generate high-fidelity synthetic data.
- Practical experience with cloud environments, big-data tools, microservices orchestration, and scalable inference engines.
- Ability to build Golden Datasets, LLM-as-a-Judge frameworks, and empirical alignment and quality evaluations.
- Proficiency in Python, PyTorch, efficient GPU utilization, and scalable API and microservice architectures.
- Preferred experience with Google Cloud Platform, BigQuery, Google Kubernetes Engine, Vertex AI, and Agent Platform.
- Preferred production model-distillation experience, AI/ML publications or open-source contributions, and advanced GPU optimization knowledge including PagedAttention, Continuous Batching, acceleration kernels, and model compilation.
Tech Stack
Categories
AI ResearchML Engineering