14 points | by amulyabaral 4 hours ago ago
3 comments
I’m honestly surprised this is better benchmark wise than the text only model. I figured the addition of vision would take away from some of the text capabilities.
I believe that multimodal training increases the robustness of latent representations regardless of which modality is being processed.
One of the early results from multimodal training is that it kinda works like cross training. Training vision helps with text tasks and visa versa.
I’m honestly surprised this is better benchmark wise than the text only model. I figured the addition of vision would take away from some of the text capabilities.
I believe that multimodal training increases the robustness of latent representations regardless of which modality is being processed.
One of the early results from multimodal training is that it kinda works like cross training. Training vision helps with text tasks and visa versa.