PercepCap: Video Captioner with Structured Spatio-Temporal Perception
Video captioning requires fine-grained spatio-temporal understanding of videos, including spatial perception of where objects are located and temporal perception of when events occur. Existing MLLMs usually generate captions directly from video inputs without…
PRISM indexes and ranks — it never republishes. The full piece lives with its author on arxiv.org.
Read on arxiv.org