9 comments

  • throwaw12 13 minutes ago

    If deepseek v4 flash is beating DeepSeek V4 Pro, can we expect new V4 Pro which is on par with Opus 5 in couple weeks (even better if it beats Opus)?

  • qtalen 11 minutes ago

    Unfortunately, DeepSeek Flash still doesn’t support multimodal; otherwise, it would offer better value than GPT 5.6 SOL.

  • WithinReason an hour ago

    Already beat Luna on price/task, by about 2x:

    https://artificialanalysis.ai/models/deepseek-v4-flash?intel...

    • spwa4 an hour ago

      Maybe I'm reading that incorrectly, but it seems to me the cost is on the X-axis.

      First, your direct comparison, Deepseek V4 Flash 0731 (max effort) $0.03 (rounded up) per task @ index 50.

      OpenAI Luna:

      * high effort $0.03 (rounded down) @ index 46

      * xhigh effort $0.04 @ index 49

      * max effort $0.07 @ index 51

      So I would say a fair statement would be "OpenAI Luna between 2x and 3x the price of Deepseek Flash, what you get is 2 to 5 times faster inference"

      The cheapest OpenAI model that beats it is OpenAI Luna (max effort) $0.07 @ index 51 (if you take the rounding out it summarizes to triple the price for similar performance), but still close to 3x faster.

      And can SOMEONE please tell artificialanalysis that using dark blue for both Deepseek AND OpenAI is an especially unfortunate choice of colors, especially today?

  • monooso 2 hours ago

    404. I believe this is the correct URL:

    https://artificialanalysis.ai/models/deepseek-v4-flash

    • theanonymousone an hour ago

      Yes, sorry, I went into anti-procrastination mode after I posted. I hope someone fixes it.

  • embedding-shape an hour ago

    Is the "Output Tokens per Intelligence Index Task" data actually correct or am I reading it wrong? It says there that "Kimi K3 (Max)" would think/reason less than than deepseek-v4-flash, and a whole bunch of other models, like less than hy3 and even gpt-oss-120b, but in my experience, K3 is probably the model that thinks/reasons the longest of all of these.

    Am I just using it on tasks that makes it go on forever vs these benchmarks that are short&sweet, or something like that? I've been throwing bunch of identical prompts at different models at the same time, and when comparing hy3 and K3 I've never once had K3 reason less than hy3, as just one anecdotal data point.

  • yonisto 4 minutes ago

    Does it already know the answer to what happen at Tiananmen Square? Or still avoiding it?

  • lostmsu 2 hours ago