Granulon: Awakening Pixel-Level Visual Encoders with Adaptive Multi-Granularity Semantics for MLLM
Published in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
Granulon combines a text-conditioned granularity controller with adaptive visual token aggregation to support fine- and coarse-grained multimodal reasoning.
Recommended citation: Junyuan Mao, Qiankun Li, Linghao Meng, Zhicheng He, Xinliang Zhou, Kun Wang, Yang Liu, Yueming Jin. (2026). "Granulon: Awakening Pixel-Level Visual Encoders with Adaptive Multi-Granularity Semantics for MLLM." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Pages 26317-26327.
Download Paper
