Training-free Multi-view 4D Human Motion Reconstruction Virtual Reality System

1Carnegie Mellon University
2National Institute of Advanced Industrial Science and Technology
WACV 2026
Multi-Cali Anything Architecture

Architecture of human and environmental reconstruction: The system processes multi-view input through a state-of-the-art single-view HMR method to estimate human meshes for each view. Person identification across views identifies the same individual in multiple views. The meshes are then refined using triangulation. For human tracking and motion smoothness, we ensure continuous and coherent reconstruction throughout the video. In addition, the 3D environment is reconstructed from selected scenes and integrated seamlessly with the human reconstruction to provide a cohesive XR experience.

Abstract

Human mesh recovery offers substantial potential for detailed behavior analysis and understanding of complex human-environment interactions. In this paper, we propose a novel 4D Human Motion Reconstruction Virtual Reality System that integrates advanced 4D multi-view human mesh recovery and high-quality 3D environment reconstruction using 3D Gaussian Splatting (3DGS). Our system seamlessly combines detailed 4D human behavior capture with accurate 3D environment reconstruction, significantly extending traditional visual monitoring approaches. Visualization through an interactive Virtual Reality (VR) platform enables dynamic interaction representation using accurately reconstructed virtual environments and computer-generated (CG) avatars. Experimental results from realistic scenarios validate the effectiveness of our framework in providing immersive experiences and precise human-environment modeling, demonstrating a significant advancement in a practical human-centered representation approach. Our approach consistently outperforms existing state-of-the-art methods, achieving reductions in mesh errors of 24% in PVE and 32% in MPJPE on the CHI3D dataset, and 17% in MPJPE and 64% in translation error on the Hi4D dataset compared to other multi-view methods.

Visualization

Multi-Cali Anything Architecture

Visualization of human reconstruction (directly test on Shelf without fine-tuning).

Quantitative Results

Multi-Cali Anything Architecture

Performance comparison on the CHI3D dataset. ”Fine-tune” indicates whether the method was fine-tuned on the CHI3D training set or directly tested without additional training. The notation of ‡ denotes the adaptation and fine-tuning.

Multi-Cali Anything Architecture

Performance comparison on the Hi4D dataset. ”Fine-tune” indicates whether the method was fine-tuned on the Hi4D training set or directly tested without additional training.