DeepSeek-V4-Flash-Vision-Exp

(huggingface.co)

14 points | by amulyabaral 4 hours ago ago

3 comments

  • syntaxing 3 hours ago

    I’m honestly surprised this is better benchmark wise than the text only model. I figured the addition of vision would take away from some of the text capabilities.

    • Llamamoe 31 minutes ago

      I believe that multimodal training increases the robustness of latent representations regardless of which modality is being processed.

    • mcbuilder 3 hours ago

      One of the early results from multimodal training is that it kinda works like cross training. Training vision helps with text tasks and visa versa.