# SpatialTree **Repository Path**: ByteDance-Seed/SpatialTree ## Basic Information - **Project Name**: SpatialTree - **Description**: CVPR 2026 (Highlight); Spatial Intelligence; MLLMs - **Primary Language**: Unknown - **License**: Not specified - **Default Branch**: main - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 0 - **Created**: 2026-09-15 - **Last Updated**: 2026-09-15 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # SpatialTree: How Spatial Abilities Branch Out in MLLMs [![arXiv](https://img.shields.io/badge/arXiv-2512.20617-b31b1b.svg)](https://arxiv.org/abs/2512.20617) [![Project Page](https://img.shields.io/badge/Project-Page-green)](https://spatialtree.github.io/) [![Dataset](https://img.shields.io/badge/Dataset-HuggingFace-yellow)](https://huggingface.co/datasets/LongfeiLi/SpatialTree-Bench) **SpatialTree** is a cognitive-science-inspired hierarchy and benchmark for evaluating spatial abilities in Large Multimodal Models (MLLMs). It organizes spatial capability into four levelsβ€”perception (L1), mental mapping (L2), simulation (L3), and agentic competence (L4)β€”spanning 27 sub-abilities. ## 🌳 SpatialTree Hierarchy Distinct from previous benchmarks, our hierarchy moves from basic perception to complex agentic interactions:
## πŸ’‘ Key Findings - **Hierarchy Matters**: Higher-level skills (L2-L4) are strongly correlated, while L1 skills are largely orthogonal. - **Transfer Dynamics**: Strong cross-level transfer exists from low-level to high-level abilities, whereas transfer within L1 can be negative. - **Auto-Think**: Naive chain-of-thought/reasoning helps complex tasks but hurts intuitive perception. We propose an **auto-think** strategy that suppresses unnecessary deliberation, enabling RL to consistently improve performance across levels. ## πŸ“Š Evaluation We provide evaluation scripts integrated with [lmms-eval](https://github.com/EvolvingLMMs-Lab/lmms-eval) to test various Large Multimodal Models (LMMs) on **SpatialTreeBench**. ### Installation First, install `lmms-eval` and the necessary dependencies: ```bash git clone https://github.com/EvolvingLMMs-Lab/lmms-eval.git cd lmms-eval pip install -e . ``` ### How to Run To evaluate a model on `SpatialTreeBench` via an OpenAI-compatible interface, use the following commands. The task is registered as `spatialtreebench`. #### Evaluate GPT-4o (via OpenAI API) Configure the judge model settings for LLM-as-a-Judge: ```bash export MODEL_VERSION="gpt-4o-2024-05-13" export API_TYPE= # openai, azure, async_openai, or async_azure export OPENAI_API_KEY=xxxx export OPENAI_API_URL= # Default: https://api.openai.com/v1/chat/completions/v1 ``` You can configure these environment variables directly or modify them in the script. Then, run the evaluation: ```bash bash examples/models/openai_compatible.sh ``` ### Metrics & Results After the evaluation finishes, the results will be saved in the directory specified by `--output_path`. The hierarchical scores will also be displayed in the terminal as follows: ```bash --- SpaTreeBench Hierarchical Scores --- SpaTree β”œβ”€β”€ L1 β”œβ”€β”€ β”œβ”€β”€ Geometry β”œβ”€β”€ β”œβ”€β”€ β”œβ”€β”€ Distance β”œβ”€β”€ β”œβ”€β”€ β”œβ”€β”€ Shape β”œβ”€β”€ β”œβ”€β”€ └── Size β”œβ”€β”€ β”œβ”€β”€ Localization β”œβ”€β”€ β”œβ”€β”€ β”œβ”€β”€ 3D Detection β”œβ”€β”€ β”œβ”€β”€ └── 3D Grounding β”œβ”€β”€ β”œβ”€β”€ Motion β”œβ”€β”€ β”œβ”€β”€ β”œβ”€β”€ Allo β”œβ”€β”€ β”œβ”€β”€ └── Ego β”œβ”€β”€ β”œβ”€β”€ Orientation β”œβ”€β”€ β”œβ”€β”€ β”œβ”€β”€ Gravity β”œβ”€β”€ β”œβ”€β”€ └── Object Orientation β”œβ”€β”€ └── Relation β”œβ”€β”€ └── β”œβ”€β”€ Correspondence β”œβ”€β”€ └── └── Relative Direction β”œβ”€β”€ L2 β”œβ”€β”€ β”œβ”€β”€ Memory β”œβ”€β”€ β”œβ”€β”€ β”œβ”€β”€ Cognitive Map β”œβ”€β”€ β”œβ”€β”€ └── Memory Retrieval β”œβ”€β”€ └── Underst. β”œβ”€β”€ └── β”œβ”€β”€ Affordance β”œβ”€β”€ └── β”œβ”€β”€ Motion Understanding β”œβ”€β”€ └── β”œβ”€β”€ Perspective Taking β”œβ”€β”€ └── β”œβ”€β”€ Relation Understanding β”œβ”€β”€ └── └── Spatial Caption β”œβ”€β”€ L3 β”œβ”€β”€ β”œβ”€β”€ Caus. Reas. β”œβ”€β”€ β”œβ”€β”€ β”œβ”€β”€ Dynamics β”œβ”€β”€ β”œβ”€β”€ β”œβ”€β”€ Geometry Puzzles β”œβ”€β”€ β”œβ”€β”€ └── Relation β”œβ”€β”€ └── Seq. Plan. β”œβ”€β”€ └── β”œβ”€β”€ Operation β”œβ”€β”€ └── └── Route └── L4 └── β”œβ”€β”€ Goal-Driven Execution └── β”œβ”€β”€ β”œβ”€β”€ Agentic Navigation └── β”œβ”€β”€ └── Robotic Arm └── └── Open-world Exploration └── └── β”œβ”€β”€ Knowledge Acquisition └── └── └── Self-Goaling ---------------------------------------- ``` You can check `results.json` for aggregated scores and `samples.json` for detailed model inputs and predictions. ### Dataset Loading The script automatically handles dataset downloading from Hugging Face: [SpatialTree-Bench](https://huggingface.co/datasets/LongfeiLi/SpatialTree-Bench). If you need to run the evaluation offline, please download the dataset beforehand and modify the `dataset_path` in `lmms_eval/tasks/spatialtreebench/spatialtreebench.yaml`. --- ## πŸ–ŠοΈ Citation If you find this work helpful, please cite our paper: ```bibtex @article{xiao2025spatialtree, title={SpatialTree: How Spatial Abilities Branch Out in MLLMs}, author={Xiao, Yuxi and Li, Longfei and Yan, Shen and Liu, Xinhang and Peng, Sida and Wei, Yunchao and Zhou, Xiaowei and Kang, Bingyi}, journal={arXiv preprint arXiv:2512.20617}, year={2025} } ```