Multimodal retrieval often returns items that are either too coarse (bringing distraction) or uniformly fragmented (losing interpretation context). Canopy's core insight is to treat each retrieved item as a hierarchy and adapt the retained granularity region-by-region: keep whole parents when context matters, descend into children when finer evidence scores at least as well as their parent, and request targeted additional retrieval only when accumulated compressed evidence is judged insufficient.
Key Findings
- Learned node encoder: fine-tuned on gold evidence to produce query–region similarity scores that are comparable across parent and child nodes, enabling score-based pruning without per-node LLM calls.
- Parent-relative refinement: a visited internal node is replaced by all children whose score s(child) >= s(parent); branches stop independently, producing an evidence forest that mixes granularities within items.
- Critic-guided additional retrieval: when compressed evidence is insufficient (especially for multi-hop questions), a critic issues a targeted follow-up query; newly retrieved items are compressed before inclusion, keeping the accumulated evidence volume bounded.
- Empirical results: evaluated over NQ, HotpotQA, OTT-QA, MMQA and LVBench using a heterogeneous ~33M-item corpus; achieves higher average answer accuracy than baseline retrieval pipelines. In an unrouted Qwen3-VL-8B-Instruct pipeline, compression cut reader-input evidence tokens by 14.2–27.7% with comparable answer accuracy. Ablations show additional retrieval drives most gains on multi-hop QA.
How it works (concise)
- Representation: each retrieved item is converted into a hierarchy of regions (whole item → sections → paragraphs/rows/frames → leaves); images are single nodes, tables may be retained whole for aggregate questions.
- Scoring: a neural node encoder scores each node against the current query; scores are learned so parent–child comparisons are meaningful for traversal.
- Refinement: perform a top-down traversal; at each internal node, if any child meets s(child) >= s(parent) the algorithm replaces the parent with all qualifying children and continues recursively only on those branches; otherwise the parent is retained intact.
- Iteration: after compression, a critic inspects accumulated evidence and, if necessary, issues targeted follow-up retrievals; the pipeline repeats until the critic is satisfied or budget is exhausted.
Who it's for and tradeoffs
Great fit if you build RAG systems handling heterogeneous corpora (text, tables, images, video) and need to limit LLM/reader input tokens while preserving interpretability across mixed-granularity evidence. Also useful when multi-hop reasoning benefits from iterative retrieval guided by partial evidence.
Look elsewhere if you cannot provide or fine-tune node-level supervision (Canopy relies on gold-evidence supervision to train the node encoder), if your pipeline requires end-to-end LLM summarization for compression, or if retrieval recall is extremely low—compression cannot recover facts that were never retrieved.
Overall, Canopy offers a lightweight, modality-agnostic post-retrieval compressor that balances context preservation and token efficiency and pairs naturally with targeted additional retrieval to improve multi-hop QA accuracy.