CastleRAG

MULTIMODAL EVIDENCE RETRIEVAL

Every answer.
Backed by
a moment.

Turn hours of video, conversations, and sensor data into evidence you can find, inspect, and question.

Explore the pipeline Explore the code

Built for verifiable question answering over CASTLE 2024.

CASTLERAG / EVIDENCE MEMORY01—03
VIDEOTRANSCRIPTSSENSORS
01

Retrieve the moment

Search across words, scenes, and signals.

02

Weigh the evidence

Rerank the most relevant evidence packs.

03

Answer with context

Ground the answer in retrieved evidence.

ONE MEMORY. MULTIPLE PERSPECTIVES.
Egocentric videoSpoken languageVisual context

01 / THE APPROACH

Find the evidence.
Then form the answer.

CastleRAG builds an offline memory of the recordings, then brings the right evidence together at question time.

HYBRID RETRIEVAL01 / 03

Different signals.
A shared evidence trail.

Keyword search finds exact language. Dense retrieval finds semantic matches across transcripts and visual content. Reciprocal rank fusion brings those results into one ranked evidence list.

BM25 + OmniEmbedRank fusionEvidence list

Speech and visual-text retrieval each contribute evidence to the multimodal search.

02 / THE EVIDENCE

More than
one point of view.

A spoken phrase tells part of the story. The camera adds context. Linked sensor data adds another perspective.

01 / VISUAL

See what happened.

Egocentric video clips, frame captions, and on-screen text preserve the visual context around an event.

Video clipsCaptions & OCR
02 / LANGUAGE

Find what was said.

Normalized transcripts make conversations searchable by exact words and by meaning.

TranscriptsHybrid search
03 / CONTEXT

Connect the signals.

Auxiliary data, including gaze, heart rate, photos, and thermal recordings, can enrich the evidence memory.

Sensor dataAuxiliary media

03 / RESEARCH IN THE OPEN

Follow the reasoning.
Explore the work.

CastleRAG is a research project targeting the CASTLE Challenge at EgoVis 2026. Explore the implementation, run the pipeline, or inspect retrieved moments in the evidence dashboard.

View the repository
DATASETCASTLE 2024

Multimodal recordings of everyday activity.

MODELQwen3-VL-8B

Generation and evidence reranking.

WORKSPACEQuestion. Inspect. Refine.

A dashboard for investigating evidence across cameras.