Composed Video Retrieval
- CoVR: Learning Composed Video Retrieval from Web Video Captions [1]: defines composed video retrieval and introduces the WebVid-CoVR dataset
- CoVR-2: Automatic Data Construction for Composed Video Retrieval [2]: Q-Former, decoupled
- Composed Video Retrieval via Enriched Context and Discriminative Embeddings (ECDE) [3]: enriches the query with a detailed reference-video description, decoupled
- EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval [4]: introduces an egocentric benchmark with action-heavy edits and hard distractors
- FDCA: Learning Fine-Grained Representations through Textual Token Disentanglement in Composed Video Retrieval [5]: fine-grained fusion of query video and text through disentangled text. Query and candidate are decoupled.
- Beyond Simple Edits: Composed Video Retrieval with Dense Modifications (BSE-CoVR / Dense-CoVR) [6]: builds Dense-WebVid-CoVR with long descriptions and dense modifications, decoupled
- From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos (TF-CoVR) [7]: targets fast, fine-grained action differences , decoupled
- CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content [8]: introduces an audio-visual CoVR setting, decoupled
- OmniCVR: A Benchmark for Omni-Composed Video Retrieval with Vision, Audio, and Text [9]: benchmarks vision-centric, audio-centric, and integrated queries, decoupled
- CoVR-R: Reason-Aware Composed Video Retrieval [10]: generate reasoning based on query video and text
- MoRe: Compositional Transformation Reasoning for Composed Video Retrieval [11]: two-stage retrieval. Stage1: Pareto multi-objective candidate selection; Stage 2: use MLLM to compare two candidate videos.
- ReCoVR: Closing the Loop in Interactive Composed Video Retrieval [12]: extends CoVR to multi-round user feedback
References
[1] Liu, Yu, et al. “CoVR: Learning Composed Video Retrieval from Web Video Captions.” AAAI, 2024.
[2] CoVR-2: Automatic Data Construction for Composed Video Retrieval. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.
[3] Zhang, Y., et al. “Composed Video Retrieval via Enriched Context and Discriminative Embeddings.” CVPR, 2024.
[4] EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval. ECCV, 2024.
[5] FDCA: Learning Fine-Grained Representations through Textual Token Disentanglement in Composed Video Retrieval. ICLR, 2025.
[6] Beyond Simple Edits: Composed Video Retrieval with Dense Modifications. ICCV, 2025.
[7] From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos. NeurIPS, 2025.
[8] CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content. ICASSP, 2026.
[9] OmniCVR: A Benchmark for Omni-Composed Video Retrieval with Vision, Audio, and Text. ICLR, 2026.
[10] CoVR-R: Reason-Aware Composed Video Retrieval. CVPR, 2026.
[11] MoRe: Compositional Transformation Reasoning for Composed Video Retrieval. CVPR, 2026.
[12] ReCoVR: Closing the Loop in Interactive Composed Video Retrieval. arXiv preprint, 2026.