BlogsIBMMultimodal Reasoning and Vision Models

Multimodal Reasoning and Vision Models

Multimodal Reasoning and Vision Models

1
posts
2025

This post introduces the Granite 3.2 models, which include new reasoning capabilities for language models and the first vision model. These models enable developers to build applications that can process and reason over multiple data types, such as text and images. Examples include turning PDFs into FAQs using multimodal RAG, debugging code from screenshots, building personal stylists, creating AI research agents for image analysis, and managing budgets from images of spending history. The post also points to tutorials and a cookbook for further exploration.

2025

Five ways to use the new Granite 3.2 models

3/17/2025

Introduces Granite 3.2 language models with reasoning capabilities and Granite-Vision-3.2-2B, the first vision model. Demonstrates use cases like multimodal RAG, image-based code debugging, AI stylist, AI research agent for image analysis, and financial assistant from spending images. Provides links to tutorials and a cookbook.