ECCV 2026, MalmΓΆ

Rethinking-based Question-Aware Frame Selection for Long-Form Video QA

better evidence Answer Doubt Look again uncertain confident: answer
Answer ↓ Doubt ↓ Look again β†Ί confident: answer Β· uncertain: rethink with better evidence

A selective rethink-and-retrieve loop for long-form video QA β€” the model answers with its initial evidence and, only when uncertainty remains, resamples frames that actually serve the question. The underlying MLLM stays frozen.

Minkuk Kim*1, Suyong Yun*1, Young Tae Kim1, Jinyoung Moon2, Jinwoo Choi†1, Seong Tae Kim†1
1Kyung Hee University    2Electronics and Telecommunications Research Institute (ETRI)
*Equal contribution   †Corresponding authors
πŸ“„ Paper arXiv πŸ’» Code πŸ–ΌοΈ Poster

Overview

How can an MLLM find the few decisive frames hidden inside a long video? It cannot simply watch everything β€” and committing to one set of frames up front means committing to whatever that set happens to miss.

ReQuest treats frame selection as part of the reasoning, not a preprocessing step. The model runs a simple loop:

  1. Answer. Reason over uniformly sampled frames and produce a first answer.
  2. Doubt. Estimate uncertainty from the model's own answer tokens. Confident cases answer directly.
  3. Look again. For uncertain cases, a lightweight question-aware selector localizes informative frames across the video, and the model reasons once more β€” this time knowing what to look for.

This selective rethink-and-retrieve process improves long-form video QA across multiple benchmarks and MLLM backbones, without modifying or fine-tuning the underlying answer model.

ReQuest teaser: uniform and similarity-based selection fail to find the decisive evidence, while ReQuest re-thinks with question-aware frames and answers Thursday.
ReQuest at a glance. Asked which weekday the woman does not exercise, uniform sampling is too sparse to find the clue and similarity-based selection chases the wrong one. ReQuest notices the missing evidence, re-thinks with question-aware frames, and lands on the answer: Thursday.

See it in action

Same video, same question β€” only the evidence changes. Pick a selection strategy to see which frames it hands the MLLM, and what answer comes back.

0:00

Notice what the Similarity tab does: it over-focuses on lexical matches, so a question about "SUBWAY" closures pulls up Subway storefronts. ReQuest retrieves evidence for the question, not keywords β€” on questions whose key object is missing from the wording, it reaches 53.0 versus 48.0 for feature-similarity selection.

Qualitative comparison: uniform sampling catches only two Spider-Men and answers 2, while ReQuest's selector peaks on frames where all three appear and answers 3.
Conclusive frames, found on demand. Asked how many Spider-Men appear in the video, uniform sampling catches only two of them and answers 2; ReQuest's selector peaks land exactly on the frames where all three appear.

Method

ReQuest first performs First Thinking over uniformly sampled frames; answer-token entropy measures how confident the prediction is. Confident questions are answered directly. Uncertain ones enter a Re-thinking stage, where dense frames are scored by a lightweight question-understandable selector (13M parameters) and K informative, non-redundant frames are picked with uncertainty-guided diversity sampling before the MLLM reasons again.

The selector is trained from the MLLM's own responses: each video segment's contribution to the correct answer is isolated against a masked-input baseline, and this contribution score supervises the selector as a pseudo label β€” no human frame annotations, and no gradients ever touch the MLLM.

ReQuest pipeline: first thinking with sparse frames, uncertainty-based routing, and question-aware re-thinking with selected dense frames.
The ReQuest pipeline. First Thinking with sparse frames, uncertainty-guided routing, and question-aware key-frame selection for Re-thinking.
Selector training: segment confidence scores from clustered frame features are compared against a masked-frame baseline to produce contribution pseudo labels.
Training the selector from the MLLM's own answers. Each segment's confidence toward the correct answer is compared against a masked-frame baseline, and the resulting contribution score becomes the selector's pseudo label β€” no human frame annotations, and no gradients ever touch the MLLM.

Results

Plugged into LLaVA-Video, ReQuest improves Video-MME by +3.0, MLVU by +7.0, and LongVideoBench by +2.1, with the largest gains on medium and long videos. The same selector transfers zero-shot to LLaVA-OneVision and Qwen3-VL without any retraining β€” and even a 512-frame Qwen3-VL benefits, where fewer, better frames beat denser uniform sampling.

Benchmark comparison

Model LLM
Size
#Frames LVB MLVU
m-avg
Video-MME (w/o sub.)
Overall Short Medium Long
Video-LLaVA7B839.147.339.945.338.036.2
VideoChat27B16–44.539.548.337.033.2
ShareGPT4Video8B16–46.439.948.336.335.0
Chat-UniVi-V1.57B64––40.645.740.335.8
VideoLLaMA27B16––47.956.045.442.1
TimeSuite7B128––46.3––41.9
Frame-Voyager7B8–65.657.567.356.348.9
LongVU7B1fps–65.460.964.758.259.5
NVILA8B102457.770.164.075.062.254.8
LLoVi–––55.154.762.153.248.8
VideoTree–––60.460.667.859.954.2
LLaVA-Video†7B3258.064.762.676.259.352.2
+ ReQuest7B3260.1+2.171.7+7.065.6+3.077.0+0.864.1+4.855.8+3.6
LLaVA-OneVision†7B3256.663.158.770.356.649.2
+ ReQuest*7B3260.2+3.668.8+5.760.9+2.271.7+1.458.8+2.252.3+3.1
Qwen3-VL†8B51262.774.070.078.670.161.2
+ ReQuest*8B≀51266.3+3.676.2+2.271.1+1.180.0+1.470.8+0.762.4+1.2

Click a column header to sort. † reproduced from the official implementation in our environment. * zero-shot transfer using a selector trained with LLaVA-Video-generated supervision. +x.x = gain over the backbone.

End-to-end cost on long videos

Because rethinking is routed by uncertainty, dense observation only runs when it is needed: ReQuest ends up faster than always-on similarity-based selection while improving long-video accuracy, with a selection overhead of about 0.3 seconds.

Frame Selection Answer Model #Frames Latency Breakdown (s) Total (s) Long Acc.
Feature Ext. Selection MLLM Inf.
No Selection (Baselines)
UniformLLaVA-Video320.5–1.62.152.2
UniformQwen3-VL-8B5123.4–11.615.061.2
Key Frame Selection (Dense Observation)
SimilarityLLaVA-Video1fpsβ†’3213.11.3Γ—10⁻³1.614.754.3
ReQuestLLaVA-Video1fps→329.40.32.712.755.8
ReQuestQwen3-VL-8B1fps→≀5128.80.316.225.162.4

Average latency measured on 900 Video-MME long-video questions; latency excludes video decoding overhead. With Re-thinking Routing, dense observation is performed only for samples routed to the re-thinking stage.

Poster

ReQuest ECCV 2026 poster

BibTeX

The BibTeX entry will be added once the ECCV 2026 proceedings are published.