This project implements an Interactive 4D (Dynamic) Point Cloud Retrieval System designed to index and search spatio-temporal 3D representations (dynamic point clouds over time) using natural language commands.
The system leverages the CL4D (Contrastive Language–4D Pretraining) model to encode dynamic point clouds and text queries into a shared embedding space, allowing users to query complex 3D human actions and physical interactions.
🌐 Official Project Website & Research: Explore the complete research paper, benchmarks, and dataset on the official project page:
👉 4D Vision UoM — CL4D: Contrastive Language–4D Pretraining for Vision-Language Reasoning in Dynamic Scenes
📐 System Architecture
🎬 Demonstration Video
Below is the demonstration of the interactive query processing and retrieval visualization interface: