Jev in 25 Lines of Python

(nobodywho.ai)

75 points | by bashbjorn 2 hours ago ago

27 comments

  • sigmoid10 an hour ago

    Going directly for the logprobs is always icky when you use a chat model as base, because they are trained to write prose as output. So your "choice" tokens and thus their probabilities might get diluted in whatever else it wanted to say. If you have to do it in the same way as this post, at least add clear system instructions and a carefully worded beginning to the assistant output section of the prompt to lower the chances of it wandering off immediately.

    I've found that using structured outputs solves this problem much better. Instead of letting a model generate only "A", "B" or "C" and looking at the probs, have it directly generate "Legitimate", "Spam" or "Phishing" or any other pre-defined option from a set of multi-token sequences. Behind the scenes it boils down to something quite similar, but you're not running into the risk that the model actually wanted to say "A phishing attempt seems likely, so answer (C) is correct.", which would lead "A" to have the highest probability in the first token. You can even use a reasoning budget this way either via inherent reasoning or a free-form part preceding the remaining output structure. You can also have it assign probabilities (either in words or numbers) using more complex output structures, but I would not rely on them much more than the token logprobs (they can still be quite good though).

    • nautilus50 32 minutes ago

      +1, llama.cpp has a --grammar parameter which you can pass a BNF style grammar file to constrain generation. It can be used in Python llama.cpp wrapper

      https://til.simonwillison.net/llms/llama-cpp-python-grammars

    • _davide_ 35 minutes ago

      Agreed, it's a real issue, but it can probably be vastly reduced by having the schema in the system prompt and by giving the model an expectation of a fixed value: no decent modern would pick a prose ligament over a provided value.

      To completely squash the issue, a few cheap LoRa iterations will do the trick just fine.

      • wongarsu 20 minutes ago

        Sure, you can fix that in a couple lines. Then a couple more lines for evaluating multiple questions on the same answer in parallel. Then a couple more lines for the confidence score (which is trivial to compute from all we have, but missing regardless). Then a harness to fine-tune an existing model to perform better on this specific task, and a collection of training data to use for that

        I think we can all agree that Jev is not rocket science. It's a good idea executed well, with marketing that might have been a tad too bold

    • ainch 33 minutes ago

      In my experience as well using logprobs to try to quantify uncertainty, LLMs are a poor fit. Neural nets in general struggle with 'calibration' --- ie. if a prediction is truly 50/50, neural nets are often prone to predicting overconfidently [0].

      I ran some tests using GPT-4 to do some basic classification a couple years ago. On ambiguous options which had to be escalated to a human, the LLM would regularly output something like a 99.8% probability, compared to 99.99% for a correct answer.

      0: https://arxiv.org/pdf/1706.04599

    • _flux 36 minutes ago

      Seems like all normal english words could risk the same, so would using short but random strings be even better?

      Actually to me it sounds it could be benchmarked if this kind of effect exists in the first place.

      • sigmoid10 31 minutes ago

        Best option would be reasoning + clear system instructions + constrained output. That is, if you have to use a chat model. Which works well enough to be sure, but hey I haven't tried raising millions of dollars when I did that 3 years ago. But perhaps I was the stupid one.

  • no-name-here an hour ago

    Beyond the missing latency and compute comparisons that Heaney commenter mentioned, also nothing about its error rate compared to Jev (nor if it even always outputs in a format the app can parse, not sure how solved that is).

    But then at the end it says it’s parody. Maybe HN title should say it’s a joke.

    • est 44 minutes ago

      latency and compute comparisons highly depends on your local setup.

      you can swith to a better model for lower error rate.

      • ricardobeat 35 minutes ago

        Which massively slows down the output. Doing this with Qwen 9B already takes you into seconds per answer territory, and Jev is supposedly frontier level intelligence.

  • dhsysusbsjsi an hour ago

    Whilst I do like reading these things for technical know how, I can sympathise with the creator of jev who now presumably has to apply an order of magnitude effort to explain why the 100 smaller things done better than this add up to a much better product.

    • ramon156 6 minutes ago

      IT's the infamous "OneDrive in 10 lines of code (SFTP)"

      While technically correct, it's not the same thing

    • jpnc 44 minutes ago

      Replace 'explain' with 'sell'. Don't forget that it's a gold rush. There's no reason to sympathize with corporations in their rush for the slice of the pie.

  • shawabawa3 15 minutes ago

    strong "You can build dropbox quite trivially by getting an FTP account, mounting it locally with curlftpfs, and then using SVN or CVS on the mounted filesystem" vibes

    You have built something like jev but not jev (for starters, the output of what you've built will be absolutely worthless, the whole reason Jev is getting so much hype is because the output is good enough)

  • onion2k 37 minutes ago

    It's fast.

    If you're comparing with something, you need to state 'fast' in relative terms. Jev is definitely fast, and if this Python takes the same time to get a decision then it's also fast. If it's 100* slower than Jev though, you shouldn't be calling it 'fast', because relatively speaking it's really, really slow.

    • _davide_ 30 minutes ago

      By design it can't be significantly slower than Jev: the prompt processing (AKA PP) is exactly the same on both and will take most of the time. Then you can process every single "question" in parallel, just predicting one or two tokens (if an answer is ambiguous with a single token) per each question, again in a single batch.

      So, fast in the LLM space and comparable with Jev.

  • cupofjoakim 18 minutes ago

    I wonder if this could be a good stepping stone to write a local prompt router to optimise what model get what prompt. I.e. if the prompt is just a lookup, send it to haiku, if it's reasoning, send it to opus and if it's implementation send it to sonnet.

    • v18a 10 minutes ago

      I was thinking the same. Haven't tried it out.

  • brap 23 minutes ago

    What I don’t understand is, why would you not want “reasoning” in a classifier?

    Speed and cost are obvious reasons, but isn’t this a tradeoff?

    • ph1l337 a minute ago

      not sure if true, but if you look at laya they use BERT type models. If jev is also using a BERT-type model it is autoregressive and therefore can't reason in the way that GPT-type models can. However, you get the advantage of being able to attend in both directions.

  • ricardobeat 38 minutes ago

    Now, can you do it in <200ms for 45 questions at once, have 0% malformed output, and any kind of meaningful benchmark? We’ll wait!

    • _davide_ 14 minutes ago

      > <200ms for 45 questions at once

      Considering your own question length: ~120 characters x 45 divided by 4.1 ~= 1317 tokens.

      So question processing at 5.5k PP(around the actual PP speed of GPT5.6 Sol) it would take around ~0.24 seconds + the context processing.

      Computing the output should be around ~20ms (at 50 tok/s), computing 45 tokens in parallel.

      > have 0% malformed output

      Pretty trivial; only the allowed output is selectable :)

      So, I keep repeating myself: Jev was a low-hanging fruit all along; no one cared, and probably no one will in a few weeks?

  • heaney-555 an hour ago

    Latency and compute comparison needed.

    • thephyber 34 minutes ago

      Is benchmarking Jev still a ToS violation?

      • ricardobeat 32 minutes ago

        Was it? That would make it unusable in any corporate setting.

  • teaonly an hour ago

    The principle is this.

  • iLoveOncall 16 minutes ago

    Nothing I hate more than bullshit articles claiming X in Y lines of code, only to use libraries abstracting hundreds of thousands of lines of code.