Multimodal
10 items tagged with "multimodal"
Models5
Gemini 3.1 Flash-Lite Image
A Google Gemini 3.1 Flash-Lite image model added to OpenRouter, providing image-focused multimodal capabilities with a 65,536-token context window.
GPT-5.4 Image 2
An OpenAI multimodal model oriented around image understanding/generation workflows, listed on OpenRouter as a new GPT-5.4 image-capable offering with a large context window.
Nano Banana 2
An image generation model in the Gemini app that uses personal context and Google Photos to create more personalized images.
Gemini 3.1 Flash Live
A low-latency, live audio-capable Gemini Flash model designed for more natural, reliable real-time voice interactions across Google products.
Gemini 3.1 Flash Image (Preview)
Google's Flash-speed image generation and editing model referenced as "Nano Banana 2" and listed on OpenRouter as a Gemini 3.1 Flash Image preview.
Benchmarks5
MMMU (Massive Multi-discipline Multimodal Understanding)
A multimodal benchmark of college-level questions requiring joint reasoning over text and images such as diagrams, charts, and figures.
VQAv2 (Visual Question Answering)
A benchmark that tests whether models can answer open-ended natural-language questions about images, balanced to reduce language-only shortcuts.
MMBench (Multimodal Benchmark)
A systematic multimodal benchmark that evaluates vision-language models across many fine-grained ability dimensions using a robustness-checked multiple-choice protocol.
DocVQA (Document Visual Question Answering)
A benchmark for answering questions about document images, testing OCR, layout understanding, and reasoning over text, tables, and forms.
ChartQA
A benchmark for answering questions about charts and plots that require visual data extraction plus arithmetic and logical reasoning over the extracted values.