Research
Since founding the iMUSE Lab at Shenzhen University in 2022, our work has pursued one question with several answers: what can a network of cameras know about a large, moving scene? We study scenes that are too big, too crowded, or too badly instrumented for a single viewpoint — and we treat imperfect observations (unknown calibration, unsynchronized streams, missing labels, occlusion) as part of the problem rather than a preprocessing step.
This agenda builds on my doctoral and postdoctoral work on wide-area multi-view crowd counting at City University of Hong Kong with Prof. Antoni B. Chan (City University of Hong Kong) — the CVPR 2019 and IJCV 2022 papers below are where the ground-plane fusion idea came from.
Core programme · Multi-view crowd intelligence
All papers →Problem
Recover population, location, identity and motion over large areas from multiple cameras — despite unknown calibration, unsynchronized streams, occlusion and the cost of annotation.
What we have established
Our earlier work established ground-plane multi-view fusion for wide-area counting and showed that it survives being moved to a new scene or a new camera layout. We then removed the explicit calibration requirement and extended the formulation from global counts into 3D, and from totals to identities and positions. SynMVCrowd now gives the field a scalable benchmark for studying counting and localization jointly.
What we are pursuing now
Following people rather than detecting them — multi-view tracking with explicit view–ground interaction (CVPR 2026) — and simulating them, so a system can be tested on situations no camera has recorded (EnvSocial-Diff, ICLR 2026). Two label-cost studies are in progress in the same direction: ranking fusion models under partial labels, and choosing which views are worth labeling.
Core programme · Robust 3D scene understanding
Problem
Recover reliable 3D structure when the observations are incomplete, poorly placed, heavily occluded, or too large to process uniformly.
What we have established
Selecting views by the reconstruction error they cause makes multi-view 3D reconstruction robust to view transformation (AAAI 2025); a diffusion prior plus farthest-view selection recovers a full object from a single image (ICIG 2025); and at urban scale we introduced UrbanBIS, a large-scale benchmark for fine-grained building instance segmentation (SIGGRAPH 2023), extended with adaptive region dividing and spatially-supervised contrastive learning.
What we are pursuing now
Human-centric 3D reconstruction at scene scale, particularly under severe occlusion and cross-view ambiguity, with reconstruction and generation treated as one problem rather than two.
Emerging · Street scenes and 4D occupancy
Problem
We are extending multi-view dynamic-scene reasoning from fixed camera networks to ego-centric driving cameras, with a current focus on future occupied space and multi-view scene generation.
What we have so far
InterOCF couples 2D image evidence with 3D structure over time for camera-only 4D occupancy forecasting; StreetDiff generates multi-view consistent street scenes from structure prompts. Both are currently preprints and are listed as such.
Other applications & collaborations
Multi-view fusion transfers to targets outside crowds: Chinese white dolphins in the open sea (ACM MMAsia 2021, begun as a conservation request) and silkie chickens counted on a farm (ChinaMM 2026, an undergraduate final-year project). The 2016 hyperspectral de-fencing paper is where my publication list begins.
Research resources
Datasets, benchmarks and code released by the group. Collaborations around these resources are welcome.
- SynMVCrowd — synthetic benchmark for multi-view crowd counting and localization (IJCV 2026)
- MVTrackTrans — multi-view crowd tracking model with MVCrowdTrack and CityTrack data (CVPR 2026)
- CVCS dataset — cross-view cross-scene multi-view counting (CVPR 2021)
- UrbanBIS — large-scale urban building instance segmentation benchmark (SIGGRAPH 2023)
- EnvSocial-Diff — environment-conditioned crowd simulation (ICLR 2026)
- Weakly-MVCC · VTR · view synchronization — code and data
Joining this agenda
The group →Students who join work on one of the three programmes above, usually on the part that is still open: tracking and simulation for the crowd programme, generation and reconstruction for the 3D programme, and occupancy forecasting for the emerging one. We release code, datasets and benchmarks where we can; representative resources are listed above.