NobleBlocks
Public

Multi-modal Alignment using Representation Codebook

Published in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) • Jun 1, 2022
Authors:
Jiali Duan
,
Li‐Qun Chen
,
Son N. Tran

Abstract

Aligning signals from different modalities is an important step in vision-language representation learning as it affects the performance of later stages such as cross-modality fusion. Since image and text typically reside in different regions of the feature space, directly aligning them at instance ...

Finding related papers...

Discussions

(0)

No comments yet

Be the first to share your thoughts!