Grok 4.7

(x.ai)

91 points | by meetpateltech 58 minutes ago

8 comments

  • moojacob 34 minutes ago
    Apparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same.

    Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.

    However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.

    My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.

    • jasonjmcghee 23 minutes ago
      For what it's worth - over the last few years or whatever, it seems like Anthropic benchmaxxes the least.

      That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.

    • dumberquestions 16 minutes ago
      Token price doesn't tell you much without knowing token efficiency.
      • moojacob 3 minutes ago
        Token price is good enough to tell you a couple things about how the lab is positioning the model.

        And plus, token efficiency isn't enough to tell you much if we are going to play that game. You might as well just measure how long it takes to complete a task. GPT is super token efficient but takes longer than Grok sometimes because Grok is unpopular and they have more spare capacity to serve your request because Musk impulsively splurged on a giant datacenter.

    • formerly_proven 2 minutes ago
      > My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish.

      The strength of the grokish dialect is similar to GPT, but weaker than claudish, but grokish is just closer to normal language on average.

      That being said I wouldn't pay for grok.

  • vessenes 1 minute ago
    Nice to see this release cadence increasing and some continued improvement in quality. I am guessing these models are basically still outcomes of the cursor team integrating with the massive amount of compute they now own: I’d imagine we will see significant step up improvements with grok 5 later this year as the team gets more experienced and confident with larger training deployments. Here’s hoping for another competitive frontier model!
  • ls1911 46 minutes ago
    after using cursor grok & trae.ai for several months , grok curor is highly superior results to trae.ai
  • sidgtm 11 minutes ago
    In my experience Grok especially inside Grok build is pretty solid choice, it’s a no nonsense model and stays on its course. Another surface where I truly enjoy the experience of using Grok model is Grok bot
  • kristofferR 30 minutes ago
    What's with the deceptive graph on top? Not including Astra can't have been an oversight, did the model compare poorly to it?
    • Jcampuzano2 22 minutes ago
      https://openai.com/index/our-decision-on-cursor-following-it...

      This explains why. Mentioned in another comment, but cursorbench explicitly tests with Cursor as the harness, and OpenAI doesn't allow them to use Astra in Cursor.

      • kristofferR 2 minutes ago
        That's not accurate. OpenAI doesn't allow Grok to provide Astra to Cursor customers anymore, but it doesn't ban anyone from using Astra via alternative harnesses.

        If Cursor wanted to include Astra in CursorBench nothing stops them, they could easily have spent an hour vibecoding in OpenAI API key support if it weren't convenient to not do that.

    • scottyah 16 minutes ago
      Deceptive? An extremely quick google search would answer your question. OpenAI pulled out of Cursor before they released Astra so it never got that benchmark.
    • Iolaum 26 minutes ago
      I wonder if that means that SpaceX evals show that they consider astra better than fable or that they hate Sam&co so much they don't want to show their stuff.
  • jmward01 14 minutes ago
    [flagged]
    • andsoitis 9 minutes ago
      Try it for software development.
      • jmward01 7 minutes ago
        I have even less trust in their not training on my data/credentials/everything on my computer.
  • simianwords 38 minutes ago
    I guess it’s only my opinion but having used grok for personal chat: it’s by far the worst one amongst Claude, ChatGPT and even Deepseek, Gemini etc.

    The personality is bland and it doesn’t work nearly as hard or even tries to help.

    • artemonster 25 minutes ago
      I used openrouter to send same prompt to qwen, derpseek, gemini and grok and found that grok does good research and produces less bullshit, especially when prompted to be critical of an idea
    • slowin 24 minutes ago
      This has been my experience as well. Grok will end tasks almost immediately and claim "Done!". It's definitely the laziest and most "dishonest" of all the models. The others aren't perfect, but I can't use Grok for any serious coding task.
    • Capricorn2481 20 minutes ago
      > The personality is bland

      I don't use Grok, but do you want your LLM to have a personality? "Personality" is exactly what people don't like about Claude.

  • toader 25 minutes ago
    How anyone that values democracy in the United States could support any of Elon's ventures is difficult to understand.
    • ctrlkctrls 17 minutes ago
      Judging by Elon's staggering success in all of his ventures I'd say you're out of touch.
      • chris_money202 0 minutes ago
        Think we all can agree he has had staggering successes, but they have all come from having massive capital from Paypal which wasn't anything super innovative, it just solved a convenient problem at a convenient time and was awarded handsomely. Elon has put his capital to work in various ways to become successful, not all of the ways being morally sound.
      • thereitgoes456 8 minutes ago
        He has had many failures, SolarCity and xAI and X and DOGE to name a few, but he has often bailed them out with his larger ventures.

        Even with his successes (Tesla, SpaceX) he has built them up in large part by bending levers of government to his advantage.

        • voidfunc 1 minute ago
          > Even with his successes (Tesla, SpaceX) he has built them up in large part by bending levers of government to his advantage.

          So what? Thats called being a maverick. He is very very good at executing on making money which is the point of business.

    • AtlanticThird 20 minutes ago
      Weird, that's the main reason I purchase all of Elon's products https://time.com/5936036/secret-2020-election-campaign/
      • thoman23 15 minutes ago
        Привет, fellow American!
    • jackfischer 14 minutes ago
      The public very much voted for massive administrative reform. Are you refering to DOGE, Elon Musk's influence on elections, something else?
      • estearum 13 minutes ago
        As if "the public" knows literally anything about how the US federal government is administered.

        If anything, they voted for reduced debt burden and they got the opposite. DOGE failed at pretty much every single one of the goals that the public arguably gave it a mandate for.

      • nibbleyou 9 minutes ago
        I personally don't like him using his position to spread fake news and racist propaganda
      • maelito 10 minutes ago
        • fourseventy 9 minutes ago
          Only sheep believe that was a Nazi salute
          • KyleTheDev 6 minutes ago
            Only sheep call other people sheep.

            Sheep often like to think themselves the wolf or coyote, it would seem.

    • ls612 13 minutes ago
      Hardly seems worse than supporting Dario’s antics at least vis a vis AI. There are no saints in this industry, only a panoply of flawed humans.
    • Romanulus 1 minute ago
      [dead]