4 comments

  • jerlendds a day ago

    VLM != LLM. Vision language models basically treat text tokens and image tokens the same. Post-training an LLM on images+text can improve its capabilities. Id recommend searching around the keyword VLM to find more resources on how multi-modal AI works.

    - https://huggingface.co/blog/vlms

    - https://en.wikipedia.org/wiki/Multimodal_learning

  • verdverm a day ago

    transformers were first an image understanding technique, the text processing came later, it's all about the training data and gradient descent, and allegedly attention

    • laruss5 12 hours ago

      It's actually the other way round - the Transformer architecture was introduced for text (machine translation) in "Attention Is All You Need" (2017). Vision Transformers, which apply it to images, came three years later in 2020: https://arxiv.org/abs/2010.11929

      • verdverm 6 hours ago

        right, it was not text generation per-se (completion/contemporary understanding) that came first, vision was before that, translation before that