Learning When to Think in Latent Visual Space: Reasoning Mode Adaptive Reinforcement Learning. To Answer or Not to Answer? Reliability-Oriented Reinforcement Learning for Visual Language Models. HER-Count: Learning Hyper-Exemplar Representation for Generalized Zero-Shot Object Counting. Generalizable Large Language Model Based Human Keypoint Localization. BMW: Bidirectionally Memory bank reWriting for Unsupervised Person Re-Identification. Deep Learning-Based 2D Human Pose Estimation: Current Status and Future Prospects. PolarPose: Single-stage Multi-person Pose Estimation in Polar Coordinates Joint Visual and Temporal Consistency for Unsupervised Domain Adaptive Person Re-Identification Multi-Scale Temporal Cues Learning for Video Person Re-Identification Pose-Guided Representation Learning for Person Re-Identification Global-Local Temporal Representations For Video Person Re-Identification Multi-Scale 3D Convolution Network for Video Based Person Re-Identification VP-ReID: Vehicle and Person Re-Identification System Pose-driven Deep Convolutional Model for Person Re-identification