The engine Every edition RSS Refreshed every 30 min 1m ago
PRISM The internet, refracted.

Home Research

PercepCap: Video Captioner with Structured Spatio-Temporal Perception

Signal strength 15/100

Video captioning requires fine-grained spatio-temporal understanding of videos, including spatial perception of where objects are located and temporal perception of when events occur. Existing MLLMs usually generate captions directly from video inputs without…

PRISM indexes and ranks — it never republishes. The full piece lives with its author on arxiv.org.

Read on arxiv.org

Same wavelength

Stories the engine considers adjacent to this one.

Research arXiv

Unified Video Dense Prediction from Disjoint Data

Scene understanding requires simultaneous prediction about geometry, appearance, and semantics. However, existing task-specific annotations are fragmented across incompatible, domain-specific datasets. Current unified systems circumvent this by restricting…

1 min 0 views

Research arXiv

GraphVid: Interactive Graph-Controllable Video Generation

Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires…

1 min 0 views