# SpatialTree
**Repository Path**: ByteDance-Seed/SpatialTree
## Basic Information
- **Project Name**: SpatialTree
- **Description**: CVPR 2026 (Highlight); Spatial Intelligence; MLLMs
- **Primary Language**: Unknown
- **License**: Not specified
- **Default Branch**: main
- **Homepage**: None
- **GVP Project**: No
## Statistics
- **Stars**: 0
- **Forks**: 0
- **Created**: 2026-09-15
- **Last Updated**: 2026-09-15
## Categories & Tags
**Categories**: Uncategorized
**Tags**: None
## README
# SpatialTree: How Spatial Abilities Branch Out in MLLMs
[](https://arxiv.org/abs/2512.20617)
[](https://spatialtree.github.io/)
[](https://huggingface.co/datasets/LongfeiLi/SpatialTree-Bench)
**SpatialTree** is a cognitive-science-inspired hierarchy and benchmark for evaluating spatial abilities in Large Multimodal Models (MLLMs). It organizes spatial capability into four levelsβperception (L1), mental mapping (L2), simulation (L3), and agentic competence (L4)βspanning 27 sub-abilities.
## π³ SpatialTree Hierarchy
Distinct from previous benchmarks, our hierarchy moves from basic perception to complex agentic interactions:
## π‘ Key Findings
- **Hierarchy Matters**: Higher-level skills (L2-L4) are strongly correlated, while L1 skills are largely orthogonal.
- **Transfer Dynamics**: Strong cross-level transfer exists from low-level to high-level abilities, whereas transfer within L1 can be negative.
- **Auto-Think**: Naive chain-of-thought/reasoning helps complex tasks but hurts intuitive perception. We propose an **auto-think** strategy that suppresses unnecessary deliberation, enabling RL to consistently improve performance across levels.
## π Evaluation
We provide evaluation scripts integrated with [lmms-eval](https://github.com/EvolvingLMMs-Lab/lmms-eval) to test various Large Multimodal Models (LMMs) on **SpatialTreeBench**.
### Installation
First, install `lmms-eval` and the necessary dependencies:
```bash
git clone https://github.com/EvolvingLMMs-Lab/lmms-eval.git
cd lmms-eval
pip install -e .
```
### How to Run
To evaluate a model on `SpatialTreeBench` via an OpenAI-compatible interface, use the following commands. The task is registered as `spatialtreebench`.
#### Evaluate GPT-4o (via OpenAI API)
Configure the judge model settings for LLM-as-a-Judge:
```bash
export MODEL_VERSION="gpt-4o-2024-05-13"
export API_TYPE= # openai, azure, async_openai, or async_azure
export OPENAI_API_KEY=xxxx
export OPENAI_API_URL= # Default: https://api.openai.com/v1/chat/completions/v1
```
You can configure these environment variables directly or modify them in the script. Then, run the evaluation:
```bash
bash examples/models/openai_compatible.sh
```
### Metrics & Results
After the evaluation finishes, the results will be saved in the directory specified by `--output_path`.
The hierarchical scores will also be displayed in the terminal as follows:
```bash
--- SpaTreeBench Hierarchical Scores ---
SpaTree
βββ L1
βββ βββ Geometry
βββ βββ βββ Distance
βββ βββ βββ Shape
βββ βββ βββ Size
βββ βββ Localization
βββ βββ βββ 3D Detection
βββ βββ βββ 3D Grounding
βββ βββ Motion
βββ βββ βββ Allo
βββ βββ βββ Ego
βββ βββ Orientation
βββ βββ βββ Gravity
βββ βββ βββ Object Orientation
βββ βββ Relation
βββ βββ βββ Correspondence
βββ βββ βββ Relative Direction
βββ L2
βββ βββ Memory
βββ βββ βββ Cognitive Map
βββ βββ βββ Memory Retrieval
βββ βββ Underst.
βββ βββ βββ Affordance
βββ βββ βββ Motion Understanding
βββ βββ βββ Perspective Taking
βββ βββ βββ Relation Understanding
βββ βββ βββ Spatial Caption
βββ L3
βββ βββ Caus. Reas.
βββ βββ βββ Dynamics
βββ βββ βββ Geometry Puzzles
βββ βββ βββ Relation
βββ βββ Seq. Plan.
βββ βββ βββ Operation
βββ βββ βββ Route
βββ L4
βββ βββ Goal-Driven Execution
βββ βββ βββ Agentic Navigation
βββ βββ βββ Robotic Arm
βββ βββ Open-world Exploration
βββ βββ βββ Knowledge Acquisition
βββ βββ βββ Self-Goaling
----------------------------------------
```
You can check `results.json` for aggregated scores and `samples.json` for detailed model inputs and predictions.
### Dataset Loading
The script automatically handles dataset downloading from Hugging Face: [SpatialTree-Bench](https://huggingface.co/datasets/LongfeiLi/SpatialTree-Bench).
If you need to run the evaluation offline, please download the dataset beforehand and modify the `dataset_path` in `lmms_eval/tasks/spatialtreebench/spatialtreebench.yaml`.
---
## ποΈ Citation
If you find this work helpful, please cite our paper:
```bibtex
@article{xiao2025spatialtree,
title={SpatialTree: How Spatial Abilities Branch Out in MLLMs},
author={Xiao, Yuxi and Li, Longfei and Yan, Shen and Liu, Xinhang and Peng, Sida and Wei, Yunchao and Zhou, Xiaowei and Kang, Bingyi},
journal={arXiv preprint arXiv:2512.20617},
year={2025}
}
```