Official Repository: A Comprehensive Benchmark for Logical Reasoning in MLLMs
-
Updated
Jun 17, 2025 - Python
Official Repository: A Comprehensive Benchmark for Logical Reasoning in MLLMs
SceneTeract: Probing and Improving Agent-Aware Activity Reasoning in 3D Indoor Scenes
Multi-modal and Vision Language Model Spatial Reasoning Benchmark
Understand, compare, and document model-training runs from Jupyter.
GateBench is a challenging benchmark for Vision Language Models (VLMs) that tests visual reasoning by requiring models to extract boolean algebra expressions from logic gate circuit diagrams.
To associate your repository with the vlm-benchmark topic, visit your repo's landing page and select "manage topics."