Grok 4.7 is built for the idea on your napkin.
SpaceXAI’s new model is not the best one you can buy. It is free to try, it works on its own for hours, and at two things, engineering and paperwork, it beats the expensive ones.
Written by Claude Opus 5. Curated, fact-checked and edited by Kevin Clubb.
How deep do you want to go? Pick a level and the article rewrites itself.
Three people who are not programmers.
Bryan runs an IT support desk and has been asked to squeeze forty new desk phones and a dozen Wi-Fi points into a comms cabinet nobody has looked at properly since it went in. Graham is a plumber with his own van, and for the first time customers are ringing him for heat pump quotes. Miriam owns a small shop and has a lease renewal on the desk that she does not fully understand.
None of them can code. All three have had the same thought this year: I’ve got an idea for the business, but I’d need to pay someone to build it. On Monday a new answer to that thought arrived. This issue is about whether it is worth the bother, and, more to the point, why they would pick this model over the better ones.
What landed, and how it went down.
The company is SpaceXAI. That is what Elon Musk’s AI lab xAI became after SpaceX bought it in February, so the rockets, Starlink and the Grok models now sit under one roof. The product is Grok 4.7, released on 21 September as its main model for coding and office work.
Without the jargon: you describe a job in plain English, the model goes away, sometimes for hours, writes the software, runs it, spots its own mistakes, fixes them, and comes back with a thing you can use. That is the whole of what the company claims is new this time.
It works longer on difficult tasks, checks its own work more carefully.
SpaceXAI · Introducing Grok 4.7 · 21 Sep 2026 · the vendor’s own claimIt also reads a lot at once: 500,000 tokens, roughly 375,000 words, so a whole house survey, the relevant standards and the code can all sit in front of it together.
You can try it for nothing in Grok Build, SpaceXAI’s own build-it tool. If your workplace already has GitHub Copilot, Bryan, it is turning up in there this week too.
Two days on, the reception is lukewarm. The Hacker News thread on the launch runs past 500 comments and the tone is underwhelmed: the release slipped about a fortnight, Musk said on X that a training step had gone wrong, and it landed the day before Anthropic shipped Claude Opus 5.5, which performs at Fable’s level for $4 in and $20 out, with OpenAI’s cheaper GPT-6 Sol and Luna arriving the same afternoon. Grok’s “cheapest at the top table” line lasted a day.
The complaint that comes up most from people actually running it is not about the scores. It is about manners.
It regularly seems to come up with terms and descriptions for things in its chain of reasoning and then uses these terms in its output assuming you understand what it’s talking about.
imron · Hacker News launch thread · observed 23 Sep 2026Worth knowing before you start, and easy enough to handle: tell it, in the brief, to write for someone who does not work in software. It is good at that when asked.
Read the same threads for what people are doing rather than what they think, and a pattern shows up that matters for our three. Nobody is replacing their main model with Grok. Plenty are adding it: as the second reviewer that finds bugs the others miss, as the model that gets to the point in legal research, as the one that writes plain English. That is the whole argument of this issue, made by strangers before I got to it.
Not the best on average. The best at two things.
On the independent testing, Artificial Analysis is direct about where this model sits.
Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index to bring SpaceXAI into the top 4 AI labs.
Artificial Analysis · Benchmarking Grok 4.7 · independent testingTop four is the honest summary, and fourth is the honest ranking: behind Claude Fable 5.1, GPT-6 Astra and Claude Opus 5, and that was before Tuesday, when Opus 5.5 went to the top of the same table at 58. Fable and Astra charge about five times as much; Opus 5.5 charges about twice. If price were the only reason to pick Grok, this would be a short article and already a dated one.
It is not the only reason. On SpaceXAI’s own comparison table there are two rows where Grok 4.7 beats Fable outright. Engineering with physics in it (circuits, power, heat, signal): 64.0% to Fable’s 56.4%, with Astra not entered. And legal paperwork, read and compared by the model working on its own: 19.6% to Fable’s 6.7%. Every model is poor at the second; Grok is three times less poor.
Those two are the vendor’s own numbers, so treat them as a claim rather than a finding until somebody independent runs the same tests. Two numbers that are independent point the same way: on finished documents Grok scores above Astra, and when it does not know something it invents an answer 29% of the time to Astra’s 51%, which matters a great deal when the document is a contract. Opus 5.5 and the new Sol were not on SpaceXAI’s table; Anthropic says Opus 5.5 performs at Fable’s level on most work, so assume a similar gap until someone runs the tests.
Where it is plainly worse: very long, fiddly programming jobs that a professional would have to finish anyway, and Opus 5.5 just made that gap wider. Bryan, Graham and Miriam are not doing those.
Pick the model that is best at your problem, not the one that is best on average. The whole argument, in one sentence
What Graham could actually build.
Call it the Heat Pump Checker. Graham types in a house: floor area, roughly when it was built, what insulation it has, the radiators, the current boiler, and the number on the main fuse next to the meter. Three things come back. A heat-loss figure for the coldest day the house should be designed for. A heat pump size, with a verdict on whether the existing radiators will cope at a lower flow temperature or need swapping. And the question that catches installers out: can the house’s electricity supply take it, does the network operator need telling first, and is the supply shared with next door?
That last part is electrical engineering. The first part is the same kind of physics, with the numbers pinned down by reality. It is the one area where Grok 4.7 beats Fable on the tests, and it ends in a finished document, which is its other strength. A description of a plausible build, not a case study; the model came out on Monday.
Why Graham’s customers care: the Boiler Upgrade Scheme pays £7,500 towards a heat pump (£9,000 for homes on oil or LPG until March 2027), but only through an MCS-certified installer, and MCS wants a room-by-room heat-loss calculation behind every install. Graham is not replacing that calculation. He is turning up to the first visit already knowing whether the job is a goer, and quoting on the spot instead of a week later.
The other two napkins pass the same test. Miriam drops in the old lease and the landlord’s proposal and gets back what changed, which clauses bite (rent review, break clause, who fixes the roof) and a list of questions for her solicitor, who still reads it. Legal paperwork is Grok’s second lead, and its habit of saying “I don’t know” rather than guessing is exactly what you want near a contract. Bryan gives it the kit list for the cabinet and gets back the power budget for the new phones and Wi-Fi points, how long the battery backup lasts, how hot the room will get, and a Wi-Fi channel plan. Electrical engineering again.
None of this is new as an idea. Issue 004 made the case that if you can describe it, you can build it. What changed on Monday is not the describing. It is that the describing now lands on a model that happens to be best in class at Graham’s particular sort of sum.
The idea on the napkin was never the hard part. Getting it built was. Issue 022 · 23 September 2026
Before you start.
Three things, and then go and have a go.
- Keep real customer details out while you build. SpaceXAI is a US company; read its data-processing terms before real names and addresses go anywhere near it. Build with made-up houses. Switch to real ones when it runs on your own kit.
- Check its work like a new apprentice’s. 29% is not zero. Test the checker against a house Graham already knows the answer for. The certified installer signs the calculation, the solicitor reads Miriam’s lease, and an electrician looks at Bryan’s cabinet before anything is switched on.
- Know the baggage. Grok, the chatbot, was temporarily banned in three countries in January over sexualised imagery. That is the chatbot, not the tool you would be using, but if it bothers you or your customers, that is a fair reason to pick another model.
Bryan, Graham, Miriam: the thing you were going to pay someone to build now takes about as long as getting three quotes for it, and for once the cheap option is also the best one at your particular job. Write the idea down properly, paste it in, and see what comes back.
Six lines, two sources.
No commentary. Every card is a verbatim quote from the independent testers or from a named person using the model, and links to it. Where a quote is shortened, the cut is marked.
“My favorite part of the new Groks has been how they speak in plain english.”
moojacob · Hacker News launch thread“the ultimate grandmaster of parallel tool-calls, routinely kicking off four or five at once.”
octoberfranklin · Hacker News launch thread“4.7 is definitely slower & more expensive … it just isn’t there for me”
mchusma · Hacker News, comparing against Sol and Opus“Grok Build with Grok 4.7 (xhigh) scores 56 on the Artificial Analysis Coding Agent Index, up +9 points from Grok 4.6 (xhigh).”
Artificial Analysis · Benchmarking Grok 4.7“Grok 4.7 (xhigh) has a lower AA-Omniscience Hallucination Rate than Grok 4.6 (high): 29% versus 34%.”
Artificial Analysis · Benchmarking Grok 4.7“Grok 4.7’s gains come with higher token usage … 81k output tokens per Intelligence Index task, compared with 36k for Grok 4.6”
Artificial Analysis · Benchmarking Grok 4.7How this issue got here.
Issue 004 argued that if you can describe it, you can build it, back when that meant paying for the best model on the market. Issue 013 is the starting point for anyone who has not yet typed anything into one of these.
21 September: Grok 4.7, free to try, fourth on the independent table, first on two vendor-reported tests that happen to be exactly the sums a tradesperson needs. 22 September: Opus 5.5 and two new GPT-6 models take away the price argument and leave the physics one standing.
Watch for independent runs of EEBench and the legal benchmark. If the two vendor numbers hold up under someone else’s testing, the case in this issue holds with them. If they shrink, so does the reason to pick this model over a cheaper-per-task rival.
Three questions. I am not certain of my own answers.
-
01 · Vendor numbers
The entire “why this model” case rests on two scores SpaceXAI published about itself. Is that good enough to recommend it to a plumber?
The counter It is good enough to try for free and test against a house you know the answer for. It would not be good enough to sign a contract on. The free tier is what makes the weak evidence acceptable, and that is not nothing.
-
02 · The second model
Everyone who rates it uses it alongside something else. Is “best at your problem” just a polite way of saying you now need two subscriptions?
The counter For Graham, no: he builds once, on a free tier, and then runs a tool that costs pennies. For a team shipping software every day, yes, and they were already paying for two.
-
03 · Who carries the risk
Graham quotes on the spot using a number a model produced. When it is wrong, whose problem is it?
The counter His, which is precisely why the MCS survey still happens and the output is marked indicative. The tool is not there to replace the certificate. It is there so he stops driving to jobs that were never going to work.
You do not need the best model on average. You need the best one at your problem.
One issue a week, and the workings underneath it.
No paywall and no pitch. If the napkin is the useful part of this issue, the two below are where the argument started: the case that describing a thing is now most of building it, and the plain-English starting point for anyone who has not begun.