See what happened.
Egocentric video clips, frame captions, and on-screen text preserve the visual context around an event.
MULTIMODAL EVIDENCE RETRIEVAL
Turn hours of video, conversations, and sensor data into evidence you can find, inspect, and question.
Built for verifiable question answering over CASTLE 2024.
Retrieve the moment
Search across words, scenes, and signals.
Weigh the evidence
Rerank the most relevant evidence packs.
Answer with context
Ground the answer in retrieved evidence.
01 / THE APPROACH
CastleRAG builds an offline memory of the recordings, then brings the right evidence together at question time.
Keyword search finds exact language. Dense retrieval finds semantic matches across transcripts and visual content. Reciprocal rank fusion brings those results into one ranked evidence list.
Speech and visual-text retrieval each contribute evidence to the multimodal search.
02 / THE EVIDENCE
A spoken phrase tells part of the story. The camera adds context. Linked sensor data adds another perspective.
Egocentric video clips, frame captions, and on-screen text preserve the visual context around an event.
Normalized transcripts make conversations searchable by exact words and by meaning.
Auxiliary data, including gaze, heart rate, photos, and thermal recordings, can enrich the evidence memory.
03 / RESEARCH IN THE OPEN
CastleRAG is a research project targeting the CASTLE Challenge at EgoVis 2026. Explore the implementation, run the pipeline, or inspect retrieved moments in the evidence dashboard.
View the repositoryMultimodal recordings of everyday activity.
Generation and evidence reranking.
A dashboard for investigating evidence across cameras.