
Vision Language Model Engineer
EchoTwin AI1 year ago
San Francisco, CA, USAMid Level
Responsibilities
- Design and implement state-of-the-art vision-language models using deep learning frameworks
- Develop and fine-tune models for image captioning, visual question answering, text-to-image generation, and other multimodal tasks
- Collaborate with data scientists and software engineers to integrate models into production systems
- Optimize model accuracy, latency, and scalability for real-world applications
- Conduct experiments, evaluate model performance, and iterate on architectures and training pipelines
- Track vision-language model research and incorporate relevant advancements
- Contribute to preprocessing, augmentation, and annotation pipelines for multimodal datasets
- Document model development processes and present findings to technical and non-technical stakeholders
Requirements
- Bachelor’s, Master’s, or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, or a related field, or equivalent experience
- At least 3 years of machine learning experience focused on vision-language models or multimodal AI
- Hands-on experience with PyTorch or TensorFlow
- Proven experience building and deploying computer vision and/or natural language processing models
- Proficiency in Python and relevant machine learning libraries such as Hugging Face, OpenCV, and Transformers
- Experience with large-scale model training and optimization, including distributed training and quantization
- Strong understanding of neural network architectures such as CNNs, Transformers, and CLIP
- Experience with multimodal datasets and image and text preprocessing techniques
- Familiarity with AWS, GCP, or Azure and model deployment workflows
- Strong problem-solving, collaboration, and communication skills
Benefits
- Medical, dental, and vision coverage options for employees and dependents in the US
- Flexible Spending Account and Dependent Care Flexible Spending Account
- 401(k) with 3% company matching
- Unlimited paid time off
- Profit sharing
- Learning and development opportunities from a diverse, multidisciplinary peer group
Tech Stack
Categories
About EchoTwin AI
EchoTwin transforms municipal fleet vehicles into autonomous sensing networks that power the Physical AI Operating System for Real-Time Urban Intelligence. By combining proprietary AI, continuous sensing, and contextual reasoning, EchoTwin transforms real-world observations into actionable intelligence and autonomous workflows, enabling governments to understand, predict, and optimize the built environment at city scale. Our mission is to build the world's first urban world model—a continuously updated digital representation of every road, asset, and piece of infrastructure that powers the next generation of city operations.