This website uses cookies

Read our Privacy policy and Terms of use for more information.


In partnership with

Hey there! 👋

Welcome back to SavvyMonk, your one-stop for AI and tech news that actually matters.

Japan's Sakana AI has launched a product that promises the power of Anthropic's restricted Mythos models without the export-control headache, and the pitch went viral within hours. The reality underneath the slogan is more interesting than the slogan itself.

Let's get into it.

The ones showing up in LLMs convert 3× better than Google

They optimized for LLMs, not just Google.

FAQs. Comparison pages. Transparent pricing. LinkedIn presence. These aren't vanity plays. They're what gets you cited in ChatGPT, Gemini, and Claude when your buyers are researching, your investors are looking, and your future hires are deciding where to work.

Download the free AEO Playbook for Startups from HubSpot and get the exact checklist. Five minutes to read.

TODAY'S DEEP DIVE

Sakana's Orchestrator Calls Other Companies' Models to Reach Frontier Benchmarks

On 22 June 2026, the Tokyo research lab Sakana AI released Fugu, a system that behaves like one model but works like a committee.

You send a request to a single endpoint that speaks the same language as the OpenAI API, and behind it a small model decides whether to answer the request itself or hand parts of it to a pool of other companies' models before stitching the results back together.

The lab shipped two versions on the same day, a faster everyday Fugu tuned for low latency and a heavier Fugu Ultra built for long, multi-step work like cybersecurity assessments and patent searches.

A Small Model Whose Only Job Is Picking Other Models

What makes the design unusual is that the coordinator is itself a trained language model of roughly seven billion parameters rather than a set of hand-written rules, and it can even call copies of itself when a problem needs another layer of checking.

The approach grows out of two papers the lab presented at ICLR 2026, Trinity and Conductor, both of which study how a model can learn to assign work across a team instead of doing everything alone. The pitch is that the scarce skill is no longer training the biggest model, it is knowing which model to reach for and when.

The Timing Is the Whole Pitch

Fugu did not arrive in a vacuum. On 12 June 2026, Anthropic pulled public access to its two most capable models, Claude Mythos 5 and Claude Fable 5, after a United States government export-control order.

Ten days later Sakana launched a product whose central promise is frontier capability without that exposure. David Ha, the lab's chief executive and a former Google Brain researcher, framed the release on X as proof that a well-coordinated pool of swappable models can stand in for restricted ones, and argued that leaning on any single company for national infrastructure or finance is now a material risk rather than a hypothetical one.

The argument lands because the event it points to genuinely happened, and any team that had built its workflow around Mythos or Fable woke up on 12 June with a problem.

The Benchmark Claim Does Not Survive Its Own Chart

Here the marketing starts running ahead of the evidence. Every score Sakana published is self-reported, measured by the company against figures other vendors quoted for their own models, which is not the same as an independent test. And even taken at face value, the headline comparison wobbles.

On the demanding SWE-Bench Pro coding test, Sakana puts Fugu Ultra at 73.7, ahead of Claude Opus 4.8 at 69.2 and GPT-5.5 at 58.6, which is a real result against the models people can actually buy. But the same chart lists Fable at 80.0, comfortably above Fugu. The product that claims to match Fable is beaten by Fable on the company's own slide.

Early independent testing has been cooler still, with Wharton professor Ethan Mollick reporting runs that stretched past half an hour and output that fell short of Fable in practice, and the precise figures quoted across coverage have not always agreed with one another.

Fugu Cannot Call the Models It Measures Itself Against

This is the detail the pitch quietly skips. Mythos and Fable are not in Fugu's pool, because they are not publicly accessible, which is the entire reason the export-control story exists in the first place.

So when Sakana says Fugu stands shoulder to shoulder with Mythos, it does not mean Fugu reaches into those models and borrows their strength. It means a coordinated team of other, available models scored close to them on a chart. Whatever Fugu routes to, it is not the restricted systems, and a buyer hoping the workaround quietly smuggles Mythos-grade output into their stack has misread what is on offer.

Routing Around a Rule Is Not Sovereignty

The sovereignty language deserves the same scrutiny. Fugu does not remove dependence on outside model makers, it spreads that dependence across several of them and reroutes when one goes dark. That is more resilient than betting everything on a single provider, but if several leading labs restricted access at the same moment, Fugu's options would shrink with them.

The routing is also deliberately opaque, since Sakana treats the question of which model handles which request as proprietary, so a user cannot see what answered them or why. The base model lets teams exclude specific providers for compliance, yet Fugu Ultra runs a fixed pool with no opt-out, and the whole service is unavailable in the European Union and the wider EEA at launch.

A product sold partly as a way around one government's export controls also invites the obvious question of how cleanly it sits with the next rule that lands.

Why It Still Matters Anyway

None of this makes the launch noise. The idea that the next gain comes from coordinating models rather than building a bigger one is a serious bet, it rests on real published research, and roughly five hundred beta users put it through long real-world tasks before launch.

Pricing starts at five dollars per million input tokens, which reads as competitive, though Sakana has said little about how much extra a single query costs once the orchestrator fans it out across several models and waits for each to finish. For anyone tired of rebuilding their stack every time a provider changes the rules, a single endpoint that absorbs that churn is a genuinely useful thing to be able to buy.

The Bottom Line

Sakana has built something real and timed it perfectly, and the orchestration idea behind it is one of the more interesting bets in the field. But the promise of Mythos-level power without export-control risk is marketing dressed two sizes too big, because Fugu cannot call Mythos, its own chart shows Fable beating it, and rerouting your dependence across several vendors is not the same as escaping it. Treat it as a smart new option to test against your own workload, not as a backdoor to the models that were locked away.

AI PROMPT OF THE DAY

Category: Vendor Claim Stress-Test

"Take this AI vendor's benchmark claim [paste the claim or chart] and stress-test it for me. Tell me which scores are self-reported versus independently verified, whether the comparison includes any model the product cannot actually access, the single best-case condition the headline number depends on, and the one question I should email the vendor to expose the gap. Give me one tight sentence for each point."

ONE LAST THING

The most quietly radical line in Sakana's launch is the claim that choosing the right model matters more than owning the biggest one. There is something to that, and it is worth sitting with as labs keep locking their best work behind regulation. But a coordinator is only ever as strong as the models it is allowed to call, so the moment to test a sovereignty promise is the moment several doors close at once. Hit reply, I read every response.

See you in the next one.

— Vivek

P.S. Know someone who builds on AI APIs and worries about the rug being pulled? Forward this to the most provider-paranoid developer you know. They can subscribe at https://savvymonk.beehiiv.com/

Reply

Avatar

or to participate

Keep Reading