Qi Zhang 张琦 iMUSE Lab · SZU

Research

Since founding the iMUSE Lab at Shenzhen University in 2022, our work has pursued one question with several answers: what can a network of cameras know about a large, moving scene? We study scenes that are too big, too crowded, or too badly instrumented for a single viewpoint — and we treat imperfect observations (unknown calibration, unsynchronized streams, missing labels, occlusion) as part of the problem rather than a preprocessing step.

This agenda builds on my doctoral and postdoctoral work on wide-area multi-view crowd counting at City University of Hong Kong with Prof. Antoni B. Chan (City University of Hong Kong) — the CVPR 2019 and IJCV 2022 papers below are where the ground-plane fusion idea came from.

Core programme · Multi-view crowd intelligence

All papers →

Problem

Recover population, location, identity and motion over large areas from multiple cameras — despite unknown calibration, unsynchronized streams, occlusion and the cost of annotation.

What we have established

Our earlier work established ground-plane multi-view fusion for wide-area counting and showed that it survives being moved to a new scene or a new camera layout. We then removed the explicit calibration requirement and extended the formulation from global counts into 3D, and from totals to identities and positions. SynMVCrowd now gives the field a scalable benchmark for studying counting and localization jointly.

What we are pursuing now

Following people rather than detecting them — multi-view tracking with explicit view–ground interaction (CVPR 2026) — and simulating them, so a system can be tested on situations no camera has recorded (EnvSocial-Diff, ICLR 2026). Two label-cost studies are in progress in the same direction: ranking fusion models under partial labels, and choosing which views are worth labeling.

Figure from Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World Scenes
Qi Zhang, Jixuan Chen, Kaiyi Zhang, Xinquan Yu, Antoni B. Chan and Hui Huang
CVPR 2026pp. 13626-13635PDFarXivCode
Figure from Mahalanobis Distance-based Multi-view Optimal Transport for Multi-view Crowd Localization
Qi Zhang, Kaiyi Zhang, Antoni B. Chan and Hui Huang*
ECCV 202415 citationsPDFarXivProject

All 16 papers in this line →

Core programme · Robust 3D scene understanding

Problem

Recover reliable 3D structure when the observations are incomplete, poorly placed, heavily occluded, or too large to process uniformly.

What we have established

Selecting views by the reconstruction error they cause makes multi-view 3D reconstruction robust to view transformation (AAAI 2025); a diffusion prior plus farthest-view selection recovers a full object from a single image (ICIG 2025); and at urban scale we introduced UrbanBIS, a large-scale benchmark for fine-grained building instance segmentation (SIGGRAPH 2023), extended with adaptive region dividing and spatially-supervised contrastive learning.

What we are pursuing now

Human-centric 3D reconstruction at scene scale, particularly under severe occlusion and cross-view ambiguity, with reconstruction and generation treated as one problem rather than two.

Figure from Multiview Multi-Person Human Mesh Recovery Under Large Scenes with Occlusions
Qi Zhang, Tao Yu, Jiechao He, Antoni B. Chan and Hui Huang
arXiv 2026PreprintPDFarXiv
Figure from UrbanBIS: a Large-scale Benchmark for Fine-grained Urban Building Instance Segmentation
Guoqing Yang, Fuyou Xue, Qi Zhang, Ke Xie, Chi-Wing Fu and Hui Huang*
SIGGRAPH 202379 citationsProject

All 5 papers in this line →

Emerging · Street scenes and 4D occupancy

Problem

We are extending multi-view dynamic-scene reasoning from fixed camera networks to ego-centric driving cameras, with a current focus on future occupied space and multi-view scene generation.

What we have so far

InterOCF couples 2D image evidence with 3D structure over time for camera-only 4D occupancy forecasting; StreetDiff generates multi-view consistent street scenes from structure prompts. Both are currently preprints and are listed as such.

Figure from InterOCF: Spatio-Temporal 2D-3D Interaction for Camera-Only 4D Occupancy Forecasting
Qi Zhang, Xinquan Yu, Kaiyi Zhang and Hui Huang
arXiv 2026PreprintPDFarXiv

All 2 papers in this line →

Other applications & collaborations

Multi-view fusion transfers to targets outside crowds: Chinese white dolphins in the open sea (ACM MMAsia 2021, begun as a conservation request) and silkie chickens counted on a farm (ChinaMM 2026, an undergraduate final-year project). The 2016 hyperspectral de-fencing paper is where my publication list begins.

Figure from TP-MVCC: Tri-plane Multi-view Fusion Model for Silkie Chicken Counting
Sirui Chen, Yuhong Feng, Yifeng Wang, Jianghai Liao and Qi Zhang*
ChinaMM 2026AcceptedUndergraduate final-year projectarXiv
Figure from Chinese White Dolphin Detection in the Wild
Hao Zhang, Qi Zhang*, Anh Nguyen Phuong, Victor Lee and Antoni B. Chan
ACM MMAsia 20216 citationsPublisherProject
Figure from Image de-fencing with hyperspectral camera
Qi Zhang, Yuan Yuan and Xiaoqiang Lu
CITS 20164 citationsPublisher

All 4 papers in this line →

Research resources

Datasets, benchmarks and code released by the group. Collaborations around these resources are welcome.

Joining this agenda

The group →

Students who join work on one of the three programmes above, usually on the part that is still open: tracking and simulation for the crowd programme, generation and reconstruction for the 3D programme, and occupancy forecasting for the emerging one. We release code, datasets and benchmarks where we can; representative resources are listed above.