11 months ago
Responsibilities
- Train and evaluate Vision-Language Models specialized for motion understanding in autonomous-driving and robotics datasets.
- Design and scale GPU-accelerated pipelines for multimodal training, fine-tuning, and inference across video, language, and sensor metadata.
- Build agentic evaluation frameworks for spatiotemporal reasoning, localization accuracy, and narrative consistency.
- Develop and productionize model-driven data curation loops that generate and refine datasets.
- Turn research breakthroughs into robust APIs and SDKs used by enterprise customers.
- Publish high-impact research while shipping features for customers.
Requirements
- Strong proficiency in Python, PyTorch, and large-scale machine-learning workflows.
- Research experience in foundation models, Vision-Language Models, or multimodal learning; publications or patents are a plus.
- Ability to run experiments end-to-end and iterate quickly and autonomously.
- Experience training or fine-tuning models on video or sensor data.
- Understanding of retrieval systems, embeddings, and GPU optimization.
- Contributions to open-source machine-learning frameworks such as DeepSpeed or Hugging Face are a plus.
- Experience with vector databases, distributed training, or machine-learning orchestration systems such as Ray, Kubeflow, or MLflow is a plus.
- Prior exposure to autonomous-driving or robotics datasets is a plus.
Categories
AI ResearchML Engineering
About Pear VC
Pear VC is an early-stage venture capital partnership that invests at pre-seed and seed, pairing capital with hands-on support, mentorship, and programs like its PearX accelerator for founding teams. Founded in 2013 and headquartered in Menlo Park, California, it backs software and technology startups across categories. Its portfolio includes companies that reached the public markets, such as DoorDash and Guardant Health.
