Official Repository: A Comprehensive Benchmark for Logical Reasoning in MLLMs
-
Updated
Jun 17, 2025 - Python
Official Repository: A Comprehensive Benchmark for Logical Reasoning in MLLMs
SceneTeract: Probing and Improving Agent-Aware Activity Reasoning in 3D Indoor Scenes
Multi-modal and Vision Language Model Spatial Reasoning Benchmark
[AAMAS 2026] OWLViz: An Open-World Benchmark for Visual Question Answering
Understand, compare, and document model-training runs from Jupyter.
GateBench is a challenging benchmark for Vision Language Models (VLMs) that tests visual reasoning by requiring models to extract boolean algebra expressions from logic gate circuit diagrams.
Project page for MindCube: can VLMs build spatial mental models of unseen space from limited views? (arXiv 2506.21458)
To associate your repository with the vlm-benchmark topic, visit your repo's landing page and select "manage topics."