publications

2026

  1. SpatialMind: Spatially Aware On-Device Embodied AI via Viewpoint Integration
    Haoming Wang, Qiyao Xue, Weichen Liu, and 3 more authors
    In Proceedings of the Annual International Conference on Mobile Computing and Networking, 2026
  2. KGRxn-LLM: Knowledge Graph Enhanced Large Language Models for Molecular Reaction Reasoning
    Weichen Liu, Qiyao Xue, Yuyang Wu, and 2 more authors
    In Proceedings of the Workshop on Biomedical Natural Language Processing (BioNLP), 2026
  3. CLORE: Content-Level Optimization for Reasoning Efficiency
    Yuyang Wu, Qiyao Xue, Guanxing Lu, and 4 more authors
    arXiv preprint arXiv:2605.22211, 2026
  4. Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning
    Yuyang Wu, Yue Huang, Shuaike Shen, and 8 more authors
    arXiv preprint arXiv:2605.07251, 2026
  5. MosaicThinker: On-Device Visual Spatial Reasoning for Embodied AI via Iterative Construction of Space Representation
    Haoming Wang, Qiyao Xue, Weichen Liu, and 1 more author
    arXiv preprint arXiv:2602.07082, 2026
  6. Reasoning Path and Latent State Analysis for Multi-view Visual Spatial Reasoning: A Cognitive Science Perspective
    Qiyao Xue, Weichen Liu, Shiqi Wang, and 3 more authors
    In Proceedings of the European Conference on Computer Vision, 2026
  7. InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity
    Haoming Wang, Qiyao Xue, and Wei Gao
    In Proceedings of the Computer Vision and Pattern Recognition Conference, 2026
  8. AAAI
    mmbert.png
    MMBERT: Scaled Mixture-of-Experts Multimodal BERT for Robust Chinese Hate Speech Detection under Cloaking Perturbations
    Qiyao Xue, Yuchen Dou, Ryan Shi, and 2 more authors
    In Annual AAAI Conference on Artificial Intelligence, 2026

2025

  1. Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
    Weichen Liu, Qiyao Xue, Haoming Wang, and 3 more authors
    arXiv preprint arXiv:2511.15722, 2025
  2. ProGait: A Multi-Purpose Video Dataset and Benchmark for Transfemoral Prosthesis Users
    Xiangyu Yin, Boyuan Yang, Weichen Liu, and 4 more authors
    In Proceedings of the International Conference on Computer Vision, 2025
  3. Phyt2v: Llm-guided iterative self-refinement for physics-grounded text-to-video generation
    Qiyao Xue, Xiangyu Yin, Boyuan Yang, and 1 more author
    In Proceedings of the Computer Vision and Pattern Recognition Conference, 2025