Gigatoken: Fastest Tokenizer

(twitter.com)

21 points | by convexstrictly 11 hours ago ago

5 comments

  • benj111 an hour ago

    I kind of assumed the model would process the text 'directly', from what I understand, wouldn't this be biasing the input based on how you tokenise as it's lossy?

    I assume this tradeoff is purely for speed/compression. Or am I missing what's going on here?

    • convexstrictly 18 minutes ago

      Tokenization is done on the CPU. Models never see the raw characters. That's why you get trick questions like the number of r's in strawberry.

      There are many research papers on models using characters directly. One challenge is that effective context length is smaller.

  • convexstrictly 11 hours ago
    • Alifatisk an hour ago

      They even added a section ”AI Use Disclosure”, beautiful!

  • doosdom 3 hours ago

    [dead]