Hey there! 👋
Welcome back to SavvyMonk, your daily dose of AI and tech news that actually matters.
Meta's superintelligence chief stood in front of staff and told them the company had finally pulled level with OpenAI. He would not say on which tests, which is the part worth holding onto.
Let's get into it.
TODAY'S DEEP DIVE
Why Meta Says Watermelon Matches GPT-5.5, And Why It's Hard To Verify
Meta's next flagship model, codenamed Watermelon, has reached the performance level of OpenAI's GPT-5.5 on closely followed benchmarks, according to remarks that Alexandr Wang made to employees at an internal town hall.
Wang runs Meta Superintelligence Labs, the division Mark Zuckerberg rebuilt around him after hiring him away from Scale AI, and his message to staff was that the money and the talent are starting to show up in the results.
That is the headline Meta wants. The problem sits one layer down, in everything the claim leaves unsaid.
What Wang Actually Told Staff
The account comes from two people familiar with the meeting, and the specific benchmarks Wang cited were never disclosed. He told the room that Watermelon is still in training and that it uses an order of magnitude more compute than its predecessor, a model internally codenamed Avocado and released publicly as Muse Spark in April 2026.
Ten times the compute is a heavy figure to attach to a model that has not finished cooking, and it frames the parity claim as something Meta is buying rather than something it has banked.
Muse Spark is the reason the boast lands with a thud for anyone who followed the last release. It performed respectably on standard tests when it arrived, but it did not top OpenAI or Anthropic, and Meta was careful at the time not to call it a frontier model. So the company is now asking the market to believe that the follow-up has closed a gap the previous model could not, on evidence that lives entirely inside Meta's own walls.
The Benchmarks Nobody Can See
An internal, single-sourced benchmark claim is not the same thing as a published evaluation, and the distinction matters more than usual here. There is no model card, no reproducible test suite, and no independent run.
Meta declined to comment, and OpenAI did not respond when asked. Until Watermelon ships and someone outside the building can measure it, caught up to GPT-5.5 is an assertion made by the people with the strongest incentive to make it.
The target is also moving faster than the claim. OpenAI released GPT-5.5 in April 2026 and followed it in late June 2026 with GPT-5.6, a more capable model that has not been released widely, reportedly at the request of the United States government.
Matching a model that is already one generation behind the leader's newest system is a real achievement and a limited one at the same time, and Meta chose to measure itself against the version it could reach.
The Coding Gap With Claude Opus
Wang carried the same message into public view. In a post on X, he said the next Muse Spark update is coming soon with large improvements in coding and agentic capabilities, aimed at being more competitive with leading models across the board. That is the general framing, and it is worth keeping separate from what came next.
When a user asked directly when Meta would ship a coding model on par with Anthropic's Claude Opus, Wang replied that it would happen pretty soon, and added that users would like what the company has cooking. He did not say Muse Spark would match Opus, and he did not put a date on it.
The coding push and the Opus comparison are two threads Wang let touch in the same thread of replies, and the honest reading is that Meta knows exactly which rival it is being measured against on the work that developers care about most.
What Meta Is Spending To Get Here
The compute story is the one part of this with hard numbers behind it. Meta has told investors it expects to spend between 125 billion and 145 billion dollars in 2026 on chips, data centres and related infrastructure, raised from an earlier range of 115 billion to 135 billion.
Zuckerberg has paired that with a hiring campaign that offered elite researchers pay packages running into the hundreds of millions each, and Wang now oversees a small internal research unit that Meta built to chase the frontier directly.
Ten times the compute on Watermelon is the brute-force answer to a gap that talent and tuning alone had not closed. It is also a bet that scale buys parity, and the benchmark claim is the first piece of evidence Meta has offered that the bet is working. The evidence would carry far more weight if anyone outside the company could check it.
Why This Matters
Strip away the codename and the town hall theatre and the real signal is the trajectory, not the score. Meta has spent a year and an enormous sum trying to convert money into model quality, and for the first time its own leadership is willing to say, in a room full of employees, that the conversion is happening.
If Watermelon ships and holds up to outside testing, it changes the shape of the race from three serious labs to four. If it ships and falls short, this town hall becomes a cautionary tale about confusing compute with capability.
The Bottom Line
Meta is asking you to trust a benchmark you cannot see, from the one party that gains the most if you believe it. The compute figures are real and the spending is staggering, but parity with a model that is already a generation old, measured on undisclosed tests, is a claim to file under promising rather than proven.
Watch for the model itself, not the town hall, because the moment Watermelon meets an independent evaluation is the moment we find out whether Meta bought its way back into the race or talked its way there.
AI PROMPT OF THE DAY
Category: Competitive Analysis
"Act as a sceptical technology analyst. A company has claimed its unreleased [Product] matches a competitor's [Product] on benchmarks, but has not disclosed which tests it ran or released the product for independent testing. Give me a checklist of the exact questions I should ask before treating this claim as credible, the specific evidence that would move it from marketing to fact, and three historical examples where a similar unverified claim later proved either accurate or hollow."
ONE LAST THING
The most valuable habit in following this industry is separating what a company has done from what it has said it will do. Meta has spent the money, that part is on the record, but the parity with GPT-5.5 exists only in a room nobody outside the company sat in.
When the party making a claim is also the only party who can see the proof, the right response is patience rather than applause. The next real data point is the model, and it is not here yet.
Hit reply, I read every response.
See you tomorrow.
— Vivek
P.S. Know a developer or founder who tracks the model race and can smell a benchmark claim from a mile off? Forward this to them. They can subscribe at https://savvymonk.beehiiv.com/

