Biography

I am currently a Ph.D. student at the School of Artificial Intelligence, Northwestern Polytechnical University (NWPU), Xi'an, China, since 2023. Before that, I received my B.S. from Shaanxi University of Science and Technology, Xi'an, China, in 2021. I am advised by Prof. Shan Gao and have been fortunate to have had the opportunity to intern at ByteDance, and Tencent, and Huawei, where I gained valuable experience in both academia and industry.

My research primarily focuses on developing robust and data-efficient visual foundation models and multimodal large language models for visual perception, understanding, and generation. More recently, I have been working on two closely related directions:

1. Robust visual foundation models: improving robustness and generalization for challenging scenarios such as degraded visual inputs, non-salient targets, and remote sensing images.

2. VLM, AIGC and Agents. Core focus areas include building multimodal evaluation benchmarks for complex scenarios, constructing Agent systems for advertising creative production, and enabling controllable generation and interactive editing for AIGC tasks.

Experience

Research Intern

Tencent

May 2026 – Present

Ad Creative Intelligence Agent

Research Intern

ByteDance

Sep 2025 – Jan 2026

Unified multimodal understanding-generation modeling based on discrete diffusion models for parallel output.

Research Intern

Huawei

Jul 2024 – Apr 2025

Robust segmentation foundation models for degraded visual inputs via generative latent-space enhancement.

News

Publications

Selected publications below. For the full and up-to-date list, please visit my Google Scholar. * denotes co-first author.
FreqDirect paper figure
Guangqian Guo, Aixi Ren, Xuehui Yu, Pengxu Wei, Yong Guo, Shan Gao Frequency Director: Learnable Mixture of Frequency Experts for Unified Concealed Scene Segmentation European Conference on Computer Vision (ECCV), 2026
GleSAM paper figure
Guangqian Guo, Yong Guo, Xuehui Yu, Wenbo Li, Yaoxing Wang, Shan Gao Segment Any-Quality Images with Generative Latent Space Enhancement IEEE Computer Vision and Pattern Recognition Conference (CVPR), 2025
VNS-SAM paper figure
Guangqian Guo, Pengfei Chen, Yong Guo, Huafeng Chen, Boqiang Zhang, Shan Gao Boosting Segment Anything Model to Generalize Visually Non-Salient Scenarios IEEE Transactions on Image Processing (TIP), 2025
GleSAM+ paper figure
Guangqian Guo, Aixi Ren, Yong Guo, Xuehui Yu, Jiacheng Tian, Wenli Li, Yaoxing Wang, Shan Gao Towards Any-Quality Image Segmentation via Generative and Adaptive Latent Space Enhancement International Journal of Computer Vision (IJCV), major revision
P2P paper figure
Guangqian Guo, Dian Shao, Chenguang Zhu, Sha Meng, Xuan Wang, Shan Gao P2P: Transforming from Point Supervision to Explicit Visual Prompt for Object Detection and Segmentation International Joint Conference on Artificial Intelligence (IJCAI), 2024
HANet paper figure
Guangqian Guo, Pengfei Chen, Xuehui Yu, Zhenjun Han, Qixiang Ye, Shan Gao HANet: Save the Tiny, Save the All — Hierarchical Activation Network for Tiny Object Detection IEEE Transactions on Circuits and Systems for Video Technology (TCSVT), 2023
RPG paper figure
*Chaowei Wang, *Guangqian Guo, Chang Liu, Dian Shao, Shan Gao Effective Rotate: Learning Rotation-Robust Prototype for Aerial Object Detection IEEE Transactions on Geoscience and Remote Sensing (TGRS), 2024
Hybrid-Net paper figure
Shan Gao, *Guangqian Guo, Hanqiao Huang, C. L. Philip Chen Go Deep or Broad? Exploit Hybrid Network Architecture for Weakly Supervised Object Classification and Localization IEEE Transactions on Neural Networks and Learning Systems (TNNLS), 2023
SAM-COD paper figure
Huafeng Chen, Pengxu Wei, Guangqian Guo, Shan Gao SAM-COD: SAM-guided Unified Framework for Weakly-Supervised Camouflaged Object Detection European Conference on Computer Vision (ECCV), 2024
P-COD paper figure
Huafeng Chen, Dian Shao, Guangqian Guo, Shan Gao Just a Hint: Point-Supervised Camouflaged Object Detection European Conference on Computer Vision (ECCV), 2024