Composed Video Retrieval

  • CoVR: Learning Composed Video Retrieval from Web Video Captions [1]: defines composed video retrieval and introduces the WebVid-CoVR dataset
  • CoVR-2: Automatic Data Construction for Composed Video Retrieval [2]: Q-Former, decoupled
  • Composed Video Retrieval via Enriched Context and Discriminative Embeddings (ECDE) [3]: enriches the query with a detailed reference-video description, decoupled
  • EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval [4]: introduces an egocentric benchmark with action-heavy edits and hard distractors
  • FDCA: Learning Fine-Grained Representations through Textual Token Disentanglement in Composed Video Retrieval [5]: fine-grained fusion of query video and text through disentangled text. Query and candidate are decoupled.
  • Beyond Simple Edits: Composed Video Retrieval with Dense Modifications (BSE-CoVR / Dense-CoVR) [6]: builds Dense-WebVid-CoVR with long descriptions and dense modifications, decoupled
  • From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos (TF-CoVR) [7]: targets fast, fine-grained action differences , decoupled
  • CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content [8]: introduces an audio-visual CoVR setting, decoupled
  • OmniCVR: A Benchmark for Omni-Composed Video Retrieval with Vision, Audio, and Text [9]: benchmarks vision-centric, audio-centric, and integrated queries, decoupled
  • CoVR-R: Reason-Aware Composed Video Retrieval [10]: generate reasoning based on query video and text
  • MoRe: Compositional Transformation Reasoning for Composed Video Retrieval [11]: two-stage retrieval. Stage1: Pareto multi-objective candidate selection; Stage 2: use MLLM to compare two candidate videos.
  • ReCoVR: Closing the Loop in Interactive Composed Video Retrieval [12]: extends CoVR to multi-round user feedback

References

[1] Liu, Yu, et al. “CoVR: Learning Composed Video Retrieval from Web Video Captions.” AAAI, 2024.

[2] CoVR-2: Automatic Data Construction for Composed Video Retrieval. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

[3] Zhang, Y., et al. “Composed Video Retrieval via Enriched Context and Discriminative Embeddings.” CVPR, 2024.

[4] EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval. ECCV, 2024.

[5] FDCA: Learning Fine-Grained Representations through Textual Token Disentanglement in Composed Video Retrieval. ICLR, 2025.

[6] Beyond Simple Edits: Composed Video Retrieval with Dense Modifications. ICCV, 2025.

[7] From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos. NeurIPS, 2025.

[8] CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content. ICASSP, 2026.

[9] OmniCVR: A Benchmark for Omni-Composed Video Retrieval with Vision, Audio, and Text. ICLR, 2026.

[10] CoVR-R: Reason-Aware Composed Video Retrieval. CVPR, 2026.

[11] MoRe: Compositional Transformation Reasoning for Composed Video Retrieval. CVPR, 2026.

[12] ReCoVR: Closing the Loop in Interactive Composed Video Retrieval. arXiv preprint, 2026.