Loading...

CLIP learns from captions, not labels — two encoders, one shared space, and an N N matrix whose diagonal must win | AIWedia