Rendered at 16:56:16 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
moojacob 42 minutes ago [-]
Apparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same.
Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.
However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.
My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.
Lucasoato 4 minutes ago [-]
> I simply cannot stand Claudish
I totally agree, it’s like that as models become more intelligent, they are less understandable by most of people... but aren’t we humans doing the same?
Aperocky 2 minutes ago [-]
Not so, the best ideas are usually the simplest to elaborate. If someone comes up with a convoluted scheme that are hard to understand or be adequately explained, it's usually fraud.
When claude speak in convoluted mess, they are often going off on tangents in real work that you asked it to do, too.
jasonjmcghee 31 minutes ago [-]
For what it's worth - over the last few years or whatever, it seems like Anthropic benchmaxxes the least.
That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.
vessenes 7 minutes ago [-]
I was going to say the reverse - claude has been the less satisfying normalized by benchmark for me in the last year. Both astra and fable have their quirks, but I am 90% codex this year up from 10% last year.
atniomn 3 minutes ago [-]
I expect the next Anthropic release to finally reduce the prevalence of Claudish
dumberquestions 24 minutes ago [-]
Token price doesn't tell you much without knowing token efficiency.
user43928 4 minutes ago [-]
Their leading benchmark with cost per task shows a tough sell compared to Fable 5.1 Low and doesn't reach the performance of Fable 5.1 Medium.
How representative that is of real world usage, I don't know.
In their benchmark GPT 5.6 Sol performs suspiciously poorly compared to the former models.
11 minutes ago [-]
9 minutes ago [-]
vessenes 9 minutes ago [-]
Nice to see this release cadence increasing and some continued improvement in quality. I am guessing these models are basically still outcomes of the cursor team integrating with the massive amount of compute they now own: I’d imagine we will see significant step up improvements with grok 5 later this year as the team gets more experienced and confident with larger training deployments. Here’s hoping for another competitive frontier model!
maz1b 4 minutes ago [-]
Either way, the fact that xAI or SpaceXAI or whatever the name is, I can commend the team behind it on their rapid ascent and progress by being close and or on the frontier in several respects.
6thbit 1 minutes ago [-]
( why is the x-axis on the first chart in descending order ? )
Tsarp 6 minutes ago [-]
Waiting on simonw "Generate an SVG of a pelican riding a bicycle " benchmark to judge this model
rvz 58 seconds ago [-]
You mean pseudo-benchmark. Might as well ask an AI model to generate audio from text or ask an AI model specifically designed to generate SVGs [0] to generate videos.
Did 4.6 not have an x-high reasoning level? Why are they comparing 4.7 x-high with 4.6 high?
andsoitis 5 minutes ago [-]
Congratulations to the team!
Saline9515 6 minutes ago [-]
I tried in Omp (Oh-my-pi), and so far it's really problematic.
It will loop in thinking mode ("Let me implement those fixes: Fix 1, Fix 2, Fix 3 .... Fix 80, Fix 81"), ignore the AGENTS.md instructions, corrupt plan files, etc etc... I have 5.6 Sol as advisor/watchdog, and it blocks every turn, I never saw this. Quite a shame, 4.6 wasn't so bad.
ls1911 54 minutes ago [-]
after using cursor grok & trae.ai for several months , grok curor is highly superior results to trae.ai
sidgtm 19 minutes ago [-]
In my experience Grok especially inside Grok build is pretty solid choice, it’s a no nonsense model and stays on its course. Another surface where I truly enjoy the experience of using Grok model is Grok bot
kristofferR 38 minutes ago [-]
What's with the deceptive graph on top? Not including Astra can't have been an oversight, did the model compare poorly to it?
This explains why. Mentioned in another comment, but cursorbench explicitly tests with Cursor as the harness, and OpenAI doesn't allow them to use Astra in Cursor.
kristofferR 10 minutes ago [-]
That's not accurate. OpenAI doesn't allow Grok to provide Astra to Cursor customers anymore, but it doesn't ban anyone from using Astra via alternative harnesses.
If Cursor wanted to include Astra in CursorBench nothing would stop them, they could easily have spent half an hour vibecoding in OpenAI API key support - if it hadn't been convenient to neglect to do that.
andsoitis 6 minutes ago [-]
Even if they could do that (workaround to include Astra in CursorBench), that has no practical consequences for Cursor users and that's what I as a Cursor user (what I use for dev, though I use ChatGPT for non-dev stuff) care about.
kristofferR 4 minutes ago [-]
It would make the benchmark way better obviously, by showing how their new model compares to their competitors, the whole point of benchmarks and graphs.
scottyah 24 minutes ago [-]
Deceptive? An extremely quick google search would answer your question. OpenAI pulled out of Cursor before they released Astra so it never got that benchmark.
kristofferR 6 minutes ago [-]
Pulled out from letting them resell Astra access, that's not a limitation on running a benchmark.
Iolaum 34 minutes ago [-]
I wonder if that means that SpaceX evals show that they consider astra better than fable or that they hate Sam&co so much they don't want to show their stuff.
Its because of this. You can't use Astra in Cursor, and cursorbench uses cursor as the harness. They can't actually benchmark it using their harness hence why its not included.
babelfish 30 minutes ago [-]
They have Astra in other benchmarks lower on the page. They just don't want to show it winning
Jcampuzano2 29 minutes ago [-]
The chart is cursorbench though and they asked about the "deceptive graph"
babelfish 33 minutes ago [-]
this is exactly it.
toader 33 minutes ago [-]
[flagged]
ctrlkctrls 25 minutes ago [-]
Judging by Elon's staggering success in all of his ventures I'd say you're out of touch.
chris_money202 8 minutes ago [-]
Think we all can agree he has had staggering successes, but they have all come from having massive capital from Paypal which wasn't anything super innovative, it just solved a convenient problem at a convenient time and was awarded handsomely. Elon has put his capital to work in various ways to become successful, not all of the ways being morally sound.
thereitgoes456 16 minutes ago [-]
He has had many failures, SolarCity and xAI and X and DOGE to name a few, but he has often bailed them out with his larger ventures.
Even with his successes (Tesla, SpaceX) he has built them up in large part by bending levers of government to his advantage.
sssilver 6 minutes ago [-]
I take issue with your use of the word "bend" here.
Can you provide specific examples of where Elon has bent the levers of government?
brandonagr2 2 minutes ago [-]
What failed with X? Usage today is higher than ever
redox99 7 minutes ago [-]
xAI is the most profitable part of SpaceX by far.
voidfunc 9 minutes ago [-]
> Even with his successes (Tesla, SpaceX) he has built them up in large part by bending levers of government to his advantage.
So what? Thats called being a maverick. He is very very good at executing on making money which is the point of business.
andsoitis 8 minutes ago [-]
> He is very very good at executing on making money which is the point of business.
Also pushing technology forward.
ls612 21 minutes ago [-]
Hardly seems worse than supporting Dario’s antics at least vis a vis AI. There are no saints in this industry, only a panoply of flawed humans.
Romanulus 9 minutes ago [-]
[dead]
jackfischer 22 minutes ago [-]
The public very much voted for massive administrative reform. Are you refering to DOGE, Elon Musk's influence on elections, something else?
estearum 21 minutes ago [-]
As if "the public" knows literally anything about how the US federal government is administered.
If anything, they voted for reduced debt burden and they got the opposite. DOGE failed at pretty much every single one of the goals that the public arguably gave it a mandate for.
serbuvlad 4 minutes ago [-]
> As if "the public" knows literally anything
Ah, yes, democracy!, except for when the public is wrong.
Who decides when the public is wrong? We do! Who decides "what the public voted for"? We do! So we are the rulers? No, of course, not, this is democracy.
You want to become the decider of when the public is wrong and of what the public voted for? TYRANT! TYRANT!
nibbleyou 17 minutes ago [-]
I personally don't like him using his position to spread fake news and racist propaganda
I have even less trust in their not training on my data/credentials/everything on my computer.
solid_fuel 7 minutes ago [-]
Seriously. They already get caught uploading everyone’s private credentials once before, one would have to be a particularly gullible rube to trust grok again. Especially with musk in charge.
simianwords 46 minutes ago [-]
I guess it’s only my opinion but having used grok for personal chat: it’s by far the worst one amongst Claude, ChatGPT and even Deepseek, Gemini etc.
The personality is bland and it doesn’t work nearly as hard or even tries to help.
slowin 32 minutes ago [-]
This has been my experience as well. Grok will end tasks almost immediately and claim "Done!". It's definitely the laziest and most "dishonest" of all the models. The others aren't perfect, but I can't use Grok for any serious coding task.
artemonster 33 minutes ago [-]
I used openrouter to send same prompt to qwen, derpseek, gemini and grok and found that grok does good research and produces less bullshit, especially when prompted to be critical of an idea
Capricorn2481 28 minutes ago [-]
> The personality is bland
I don't use Grok, but do you want your LLM to have a personality? "Personality" is exactly what people don't like about Claude.
Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.
However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.
My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.
I totally agree, it’s like that as models become more intelligent, they are less understandable by most of people... but aren’t we humans doing the same?
When claude speak in convoluted mess, they are often going off on tangents in real work that you asked it to do, too.
That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.
How representative that is of real world usage, I don't know.
In their benchmark GPT 5.6 Sol performs suspiciously poorly compared to the former models.
[0] https://quiver.ai/
It will loop in thinking mode ("Let me implement those fixes: Fix 1, Fix 2, Fix 3 .... Fix 80, Fix 81"), ignore the AGENTS.md instructions, corrupt plan files, etc etc... I have 5.6 Sol as advisor/watchdog, and it blocks every turn, I never saw this. Quite a shame, 4.6 wasn't so bad.
This explains why. Mentioned in another comment, but cursorbench explicitly tests with Cursor as the harness, and OpenAI doesn't allow them to use Astra in Cursor.
If Cursor wanted to include Astra in CursorBench nothing would stop them, they could easily have spent half an hour vibecoding in OpenAI API key support - if it hadn't been convenient to neglect to do that.
Its because of this. You can't use Astra in Cursor, and cursorbench uses cursor as the harness. They can't actually benchmark it using their harness hence why its not included.
Even with his successes (Tesla, SpaceX) he has built them up in large part by bending levers of government to his advantage.
Can you provide specific examples of where Elon has bent the levers of government?
So what? Thats called being a maverick. He is very very good at executing on making money which is the point of business.
Also pushing technology forward.
If anything, they voted for reduced debt burden and they got the opposite. DOGE failed at pretty much every single one of the goals that the public arguably gave it a mandate for.
Ah, yes, democracy!, except for when the public is wrong.
Who decides when the public is wrong? We do! Who decides "what the public voted for"? We do! So we are the rulers? No, of course, not, this is democracy.
You want to become the decider of when the public is wrong and of what the public voted for? TYRANT! TYRANT!
Sheep often like to think themselves the wolf or coyote, it would seem.
Fuck, it is like the denial around Jan 6th. Those idiots we’re live streaming that shit. I watched it go down live. Now they say they weren’t violent.
We can’t have discourse when we have legit video evidence and people refuse to open their eyes and choose to deny reality
The personality is bland and it doesn’t work nearly as hard or even tries to help.
I don't use Grok, but do you want your LLM to have a personality? "Personality" is exactly what people don't like about Claude.