5 months ago
Responsibilities
- Act as the technical owner for enterprise customer vision-language model post-training engagements
- Translate customer requirements into multimodal post-training specifications and workflows
- Design and execute visual data generation, filtering, and quality assessment processes
- Develop image-text pair curation, annotation, and synthetic data generation pipelines for visual tasks
- Run supervised fine-tuning, preference alignment, and reinforcement learning workflows for vision-language models
- Design and interpret evaluations for visual understanding, grounding, OCR, document parsing, and other multimodal capabilities
- Feed evaluation learnings into core post-training pipelines and baseline model development
Requirements
- Hands-on experience with data generation and evaluation for VLM or multimodal post-training
- Experience training or fine-tuning vision-language models using supervised fine-tuning, preference alignment, and/or reinforcement learning
- Strong intuition for visual data quality, annotation design, and multimodal evaluation
- Familiarity with vision encoders, image-text architectures, and interactions between visual representations and language model backbones
- Experience with visual grounding, document understanding, OCR, or video understanding is preferred
- Experience contributing to shared or general-purpose multimodal post-training infrastructure is preferred
- Prior exposure to customer-facing or applied ML delivery environments is preferred
- Familiarity with multimodal alignment or reinforcement learning techniques beyond basic supervised fine-tuning is preferred
Benefits
- 100% employer-paid medical, dental, and vision premiums for employees and dependents
- 401(k) matching up to 4% of base pay
- Unlimited PTO and company-wide Refill Days
Categories
AI ResearchML Engineering
About Liquid AI
We build efficient general-purpose AI at every scale.
