SpaceXAI's Grok 4.6 Scores 61 on the Artificial Analysis Intelligence Index

(artificialanalysis.ai)

109 points | by wertyk an hour ago ago

56 comments

  • small_model 3 minutes ago

    SpaceXAI is the only frontier model company that had its own compute/date centres and soon chip making factory, I think they will pull ahead with cheaper tokens similar intelligence and better harness/tools. Grok build is 2-5x faster than Claude Code in my opinion.

    • bdangubic a minute ago

      file this under never gonna happen (this folder is piling up...) :)

  • ilreb a minute ago
  • satvikpendem an hour ago

    Cursor, since Grok 4.5, has had an incredible deal for frontier level models, their subscription now goes way further than OpenAI or Anthropic. Even on their lower tier plans you can use a lot tokens on their of their first party models (Grok and Composer) and not really run out comparatively. Combine them with an orchestrator and implementor type setup and it goes even further.

    • hmokiguess 33 minutes ago

      I believe they are the only western provider that has Kimi K3 on a subscription plan today as well. I would love to ditch Anthropic and be on Kimi if there were a subsidized plan like that with ZDR

      • pkaye 13 minutes ago

        GitHub Copilot does have Kimi K3.

      • homakov 13 minutes ago

        kimi k3 credits end in just a few sessions. Only Grok models allow generous use in Cursor Pro/+

      • tekwarder 9 minutes ago

        GabAI has KimiK3

    • aliljet an hour ago

      Can you explain what you mean? These days courtesy of an addictive reset game OpenAI is playing, I can't find anything with frontier intelligence that's more cost efficient...

    • nomilk 20 minutes ago

      How does Grok 4.5 compare to Opus >= 4.8 though?

      I'm willing to pay 2x for a 10% smarter model. Intelligence matters that much (because 10% smarter probably saves, on average, several hours of human time).

      • redox99 14 minutes ago

        It's a bit worse.

        I haven't tried so it's pure speculation based on benchmarks, but I'd assume Grok 4.6 is around Opus 4.8 in real world use, but clearly below Opus 5.

    • jesse_dot_id an hour ago

      Goes even further to exfiltrate your data, yeah.

      • greenavocado 43 minutes ago

        That would be Muse Spark Contributor Tier. 12-21x price reduction at the expense of your digital existence.

        • CuriouslyC 19 minutes ago

          I'd be the first model I'd reach for if I was providing a free service to AI gooners though. Serves them both right.

  • nylonstrung an hour ago

    I have never met a single human being who uses Grok for coding

    • jm4 an hour ago

      I have. He was using it due to philosophical reasons the same way many people have philosophical reasons for avoiding it. I don't know how many people are like that, but it's not exactly where you want to position your product if you're a business.

      Personally - and I know I'm not alone with this sentiment based on comments I see on this site - I wouldn't touch Grok no matter how good or cheap it is. I don't trust Elon and I don't want to give another dollar to the world's richest person who turns around and uses the money to interfere with elections. The guy I know uses it for essentially the same reason I won't use it.

      • inferniac 13 minutes ago

        >who turns around and uses the money to interfere with elections

        thats a very dumb reason considering all rich people do it, most are just not as open about it as Musk

        • MisterKent 5 minutes ago

          "Everyone's doing it" presented with no evidence is just you trying to make yourself feel better for: - using from - driving a Tesla - voting for Trump

          It's a really, really easy line to draw in the sand: don't support openly corrupt individuals.

          It's absolutely true there's money in politics. To call all money in politics equally corrupt because Bernie got a dollar to have dinner with someone vs Elon effectively directly buying votes....

          A complete lack of nuance here. And I wouldn't be surprised if corporations / super rich WANT you to think like that. The more defeatist the mentality becomes the more we just accept whatever they do next.

    • nomilk 18 minutes ago

      I had a security incident the other day and Grok was the only model that would help. Claude and GPT refused on ethical grounds and only gave general advice. In an emergency, I'd only trust Grok. However, that's the only time I used Grok for coding (since Opus 4.8 it would take a lot to get me to switch away from Anthropic)

    • busch_j an hour ago

      A bunch of SWEs at my work use it as their primary model.

      We have Claude, ChatGPT, and Cursor with essentially no cap on spend (top guy is spending over 10K a month on AI at API prices), and he hasn't had his hand slapped.

      So it's not like they are using it purely because it's cheaper.

      I think people like to use it for its speaking style, pretty solid performance, and its speed.

      • redox99 12 minutes ago

        That's pretty surprising. Idk about Grok 4.6, but Grok 4.5 was clearly below Fable, Opus 5 and GPT 5.6 Sol.

    • visopsys an hour ago

      I use. I used to be a Claude user. Since trying Grok 4.5 and especially Grok 4.6, I don't want to go back to Claude any more (I have early access to 4.6).

      Grok is 3x+ faster than Claude and I can't tell the diff in engineering work quality. As an engineer, speed is important to me.

      • bigyabai an hour ago

        For $30/month, I'd expect it to have higher usage limits than Claude Code and Codex.

        • visopsys an hour ago

          It really does. I felt like I could have spent $1000+ api token on claude for the amount of work on my $30 grok subscription.

          • drewnick an hour ago

            An hour in, I've been running four terminals full bore on my $20/mo Grok sub and I'm at 9% for the week. Codex or Claude would easily have hit 5-hour or weekly limits.

    • LeBit an hour ago

      I refuse to use that product because of the parent company.

      • wilburx3 21 minutes ago

        100%, it's an easy pass given that it is always playing catch up.

      • jryle70 17 minutes ago

        I can respect if you say you hate their guts. Everyone has their worldview. But over moral or ethical stand? You don't have any if you're using Chinese models, or fly Middle East airlines, or countless of other products. Don't delude yourself.

        • epolanski 8 minutes ago

          Chinese models are open, Grok is not.

    • peder an hour ago

      Does it matter? Why turn it into a popularity contest?

    • homakov 11 minutes ago

      once my codex/claude weekly limit was gone, i gave it a try. It was surprisingly good, not dumb in any way, and fast. I now require it as a part of 3-of-3 quorum with any codebase change.

    • kvirani an hour ago

      Folks working in US govt tend to, based on convos I've had with one such person.

      • dogmayor 7 minutes ago

        Not the best endorsement given the current US gov

    • inferniac 11 minutes ago

      it only very recently became competetive, if they proceed with improvements (and beating others on price) their share will grow

    • Recurecur 44 minutes ago

      Hi! Grok’s worked quite well for my use cases…

      It’s also a great deal!

    • treexs 25 minutes ago

      the new models are quite good, give it a shot

    • supriyo-biswas an hour ago

      I'm only being forced to use it at $WORK since some people overran their Cursor bill, so everyone gets Cursor Auto enabled by default which routes to Grok 4.5.

    • mohamedkoubaa 15 minutes ago

      I use it because it's cheap and good enough

    • sidcool an hour ago

      Hello. Nice to meet you.

    • locknitpicker 35 minutes ago

      > I have never met a single human being who uses Grok for coding

      Me too. The only people I ever saw using grok were using it by accident as they used copilot in auto mode and noticed some prompts were thrown it's way.

      I saw far more people using Mistral than grok.

    • DetroitThrow an hour ago

      I've tried it on my "let's run every model in parallel and see which finds more edge cases" type of tasks, and Grok 4.5 was really behind Opus/ChatGPT but ahead of Gemini - despite having a strong showing on benchmarks.

      That makes me really skeptical of it being GPT5.6-tier, much less Fable-tier, based on some of these benchmarks alone. But I'll test here shortly.

  • sidcool 38 minutes ago

    Grok is not the best model around, but it's decent. It gets the basic job done at a low price. I don't think it can advance frontier Math, yet.

    • leerob 26 minutes ago

      Probably can't advance frontier math yet, yeah. But please let us know other places you want to see Grok improve for future models!

      • solid_fuel 7 minutes ago

        I can’t think of any improvements you could make to grok that would get people to trust grok, Musk, or “X” with their data.

        The people behind it have repeatedly demonstrated a complete lack of morals at best, and depraved personalities at worst.

        I suppose you could start by removing the CSAM and nonconsensual pornography generation features though.

      • freejazz 3 minutes ago

        More deepfake porn creation please! Oh yeah, and for children too!

  • insane_dreamer 2 minutes ago

    Why is Grok so much cheaper than Claude or GPT?

  • pzo an hour ago

    Seems the cache read pricing almost doubled from $0.30 in Grok 4.5 to $0.50 in Grok 4.6.

    In my experience in heavy coding sessions most pricing is just cache read and cache write like 80% of my token bill.

  • petesergeant 34 minutes ago

    Interesting. Grok 4.5 is a capable model, although not quite at Fable/Sol levels. Will be interesting to see how this holds up. Musk appears to have made a savvy choice buying Cursor's data.

  • thiago_fm an hour ago

    I often wonder if there's a chance, even if minimal... that they stole the weights of the Anthropic models they run on their datacenter... or are actively destillating it.

    • connicpu an hour ago

      I think the more likely explanation is that the Cursor data they effectively acquired for $10B was extremely valuable for their training when combined with the insane number of GB300s xAI has for training.

      • winstonp an hour ago

        Cursor was 60B. The 10B number was the breakup fee if the deal fell through.

        • bpodgursky 3 minutes ago

          $60B in SPCX stock, which is sort of magic money.

    • qudat an hour ago

      > ... or are actively destillating it.

      I just assumed every model manufacturer is distilling from the frontier models. If they aren't they are definitely trying to do it.

  • TSiege 14 minutes ago

    I’m perpetually surprised people feel comfortable sending their code to an LLM that was made to call itself MechaHitler. I wouldn’t touch grok or now cursor with a 10 foot pole

    • dmix 4 minutes ago

      MechaHitler was something that existed only on X's grok chatbot, due to a one-line system prompt change they reverted after half a day.

      That's different than using Grok as a model for coding.

      • freejazz 2 minutes ago

        oh, it was only due to a "one-line system prompt change"? well okay then! Here I was thinking it was multiple lines! How foolish was I??