Sub-1-Bit LLM Compression via Latent Factorization

(github.com)

41 points | by brainless 3 hours ago ago

4 comments

  • big-chungus4 17 minutes ago

    Can this produce a useful model? So far 1 bit quants have been less useful than smaller models that use the same memory

  • badatnames 21 minutes ago

    Their paper shows this comes with huge quality loss, but that doesn't make it a negative result by any means

  • nico 19 minutes ago

    Has anyone tried this on apple silicon M1-5? Any benchmarks/comps?