> Findings derived from two curated layers: which model cards mention each benchmark, and which scores could be read verbatim from those documents. Each finding names the evidence behind it.
This is 100% AI generated but the problem is it's difficult to understand - what layers is it talking about, is it the llm model layers or what.
I am waiting with bated breath to read, “load-bearing” somewhere… The latest models are very capable, but sometimes they seem to get so deep in the details that they lose the overall plot.
What the hell is the point of this page? Can you put in a single bit of human prose explaining why it exists and what we are supposed to learn?
Good idea, would be interesting to cross-examine the benchmarks, but the page information is completely obscured by the AI slop. The benchmarks comparison and should start immediately instead of having random completely arbitrary complex headers and labels
As an overall pro-AI person... Please don't let it do the entire UI design for your site. 99% of this site is completely useless.
> What the two layers say
> Stated findings
> Findings derived from two curated layers: which model cards mention each benchmark, and which scores could be read verbatim from those documents. Each finding names the evidence behind it.
This is 100% AI generated but the problem is it's difficult to understand - what layers is it talking about, is it the llm model layers or what.
I’m getting good at reading LLM-speak (which is probably sad)
It means two “layers” of information access/validation.
Sounds like it first made a little map of model <-> benchmark, then went and filled in the score boxes.
Definitely not LLM layers
> This measures vendor attention, not benchmark quality
This text on this page is so aggressively LLM-written (Claude) I am struggling to understand what I am even looking at.
I am waiting with bated breath to read, “load-bearing” somewhere… The latest models are very capable, but sometimes they seem to get so deep in the details that they lose the overall plot.
What the hell is the point of this page? Can you put in a single bit of human prose explaining why it exists and what we are supposed to learn?
Good idea, would be interesting to cross-examine the benchmarks, but the page information is completely obscured by the AI slop. The benchmarks comparison and should start immediately instead of having random completely arbitrary complex headers and labels
Broken on mobile safari