Composed Image Retrieval

  • MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query [1]: multilingual, multi-image, and multi-condition composed retrieval; the closest work to the interleaved multi-reference query described in the PPT.
  • Pic2Word: Mapping Pictures to Words for Zero-shot Composed Image Retrieval [2]: maps a reference image to a pseudo-word in the text space and combines it with the modification text for zero-shot CIR.
  • Image Retrieval on Real-life Images with Pre-trained Vision-and-Language Models [3]: introduces CIRR, a real-image composed image retrieval benchmark with natural-language modifications.
  • Fashion IQ: A New Dataset towards Retrieving Images by Natural Language Feedback [4]: introduces FashionIQ, a fashion-domain benchmark where natural-language feedback describes changes from a reference item.
  • PinPoint: Evaluation of Composed Image Retrieval with Explicit Negatives, Multi-Image Queries, and Paraphrase Testing [5]: evaluates explicit negatives, multi-image queries, and paraphrase robustness instead of relying only on standard CIR test cases.

References

[1] MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query. NeurIPS, 2025.

[2] Saito, Kuniaki, et al. “Pic2Word: Mapping Pictures to Words for Zero-shot Composed Image Retrieval.” CVPR, 2023.

[3] Liu, Zheyuan, et al. “Image Retrieval on Real-life Images with Pre-trained Vision-and-Language Models.” ICCV, 2021.

[4] Guo, Zhaopeng, et al. “Fashion IQ: A New Dataset towards Retrieving Images by Natural Language Feedback.” CVPR, 2021.

[5] PinPoint: Evaluation of Composed Image Retrieval with Explicit Negatives, Multi-Image Queries, and Paraphrase Testing. CVPR, 2026.