Fair point on the wording — I've reworded it. And thanks for the list. AI Edge and MLC use the GPU, which I don't yet, so PocketPal is the closest comparison to mine. It was slower than PocketPal and that's fixed now.
This is very interesting. I have not tested it (yet) but I believe this is what will come next when talking about "AI will be everywhere".
The self-contained disconnected mode is very convenient, though it could be useful to have an "online access" mode to allow it to access the Internet and perform lookup.
The claim that these models "run on almost any phone" as well as the in-app indicator of performance that says it may "run comfortably" are some serious exaggeration. Tried running Qwen 0.5B on mine [1]. It took five minutes to produce "I am a large language model created by Anthropic" (lmao) and then stopped outputting anything at all. I am not expecting you to somehow make the models perform better or whatnot, I just believe that the performance claims need review.
Alright, I have since experimented with the other harnesses mentioned in this thread (AI Edge and PocketPal) and both run the same models much, much faster. Gemma 4 E2B is extremely usable, Qwen simply flies. Same device. I don't know what you are doing, but it's clearly not great...
"Download the apk from my website. Android will warn you because it did not come from the Play Store - that is normal for a direct download."
mmmh not the best presentation.
Established alternatives:
- Google AI Edge Gallery
- MLC Chat
- PocketPal AI
Fair point on the wording — I've reworded it. And thanks for the list. AI Edge and MLC use the GPU, which I don't yet, so PocketPal is the closest comparison to mine. It was slower than PocketPal and that's fixed now.
This is very interesting. I have not tested it (yet) but I believe this is what will come next when talking about "AI will be everywhere".
The self-contained disconnected mode is very convenient, though it could be useful to have an "online access" mode to allow it to access the Internet and perform lookup.
Thanks. An optional online lookup mode is a nice idea and I've noted it — offline stays the default though, that's the whole point of it.
The model is just the executing part. You can always give them tools via the harness to do various things.
The claim that these models "run on almost any phone" as well as the in-app indicator of performance that says it may "run comfortably" are some serious exaggeration. Tried running Qwen 0.5B on mine [1]. It took five minutes to produce "I am a large language model created by Anthropic" (lmao) and then stopped outputting anything at all. I am not expecting you to somehow make the models perform better or whatnot, I just believe that the performance claims need review.
[1]: https://m.gsmarena.com/xiaomi_redmi_note_13_pro-12581.php
Alright, I have since experimented with the other harnesses mentioned in this thread (AI Edge and PocketPal) and both run the same models much, much faster. Gemma 4 E2B is extremely usable, Qwen simply flies. Same device. I don't know what you are doing, but it's clearly not great...