Just before the holiday break, the AI community went wild over a paper claiming GPT-5.5 packs nearly 10 trillion parameters. Today, that same paper has been torn apart by researchers. After fixing the flaws, the real number looks more like 1.5 trillion.
In late April 2026, a paper titled Incompressible Knowledge Probes, or IKP for short, sent shockwaves through the AI world.

Paper link: https://www.alphaxiv.org/abs/2604.24827
Bojie Li, chief scientist at Pine AI, published a study claiming a brand new black-box probing method could reveal the true size of closed-source models.
The numbers exploded across social media overnight.
Think about it. If GPT-5.5 really hit 10 trillion parameters, that would make it more than five times larger than the rumored 1.8 trillion parameters of GPT-4.
Within hours, the 10 trillion figure was everywhere.

Then, just a few days later, the story flipped.
The Logic Hole: How 10 Trillion Shrunk to 1.5 Trillion
Recently, Lawrence Chan from UC Berkeley’s CHAI lab and Ben Sturgeon from UK AISI dug deep into the paper.

They found that this viral paper, which claimed to reverse-engineer the size of large models, suffered from serious logic and code errors.

After fixing these issues, GPT-5.5’s estimated parameter count dropped to about 1.5 trillion, with a 90 percent confidence interval of 256 billion to 8.3 trillion.


Where Did the Numbers Go Wrong
In the original paper, the researchers did not break down model scores by individual question. Instead, they used a flooring method. When calculating average scores across different datasets, they quietly lowered the scores of smaller models. This created a fake gap.
In plain English, when a model already knows the answer to a question, the score should be perfect. But the flooring trick made it look like the model only got partial credit.
When researchers removed the flooring and recalculated, the score gap between small and large models shrank dramatically. The original score-to-size curve was flat. Once fixed, GPT-5.5’s estimated size crashed from 9.7 trillion down to 1.5 trillion.



A 25 Percent Knowledge Gap: What the Fix Really Shows
The researchers also found that the knowledge gap between models was much smaller than the paper claimed.
The funniest part? The original author Bojie Li admitted it himself. He said the paper was written in a rush under the hype of AI coding tools, and the entire probing experiment was done in just four hours.
Lawrence Chan jokingly called the original paper vibe-coding research, a nod to the rough and ready style of AI-assisted coding.


The Real Size of Top AI Models After the Fix
In short, the core idea of IKP is not wrong. The score does relate to model size. But the way the paper turned that score into a parameter count was deeply flawed.

After the fix, the estimated sizes of top models all dropped sharply.
GPT 5.5 went from 9.7 trillion to 1.5 trillion.
Claude Opus 4.7 went from 4.0 trillion to 1.1 trillion.
DeepSeek R1, which actually has 671 billion parameters, was estimated at 424 billion to 760 billion. At least this one was in the right ballpark.

The good news is that all the corrected numbers still show a clear link between IKP scores and model size.

In other words, bigger models still score better on math. The IKP method itself is sound. It just needed better math.
Some people compare a model to a hard drive. The more space it has, the more facts it can store. But in reality, a model’s knowledge density matters far more than raw size. The idea of incompressible knowledge probes is exactly about measuring how much real knowledge a model holds, not just how big it is.
Who Holds the Real Knowledge Crown
So, does a smaller parameter count mean a weaker model? Not at all. Efficiency is what counts.
Thinking Mode: The Real Game Changer
Lawrence Chan’s latest blog post showed something even more interesting. When models enter thinking mode, their knowledge scores jump again. This proves that reasoning depth, not just memory size, is what separates good models from great ones.

GPT-5.5 Is NOT 9.7 Trillion
On April 30, Bojie Li of Pine AI responded to the criticism with a new blog post.

His core argument was that model size and IKP scores do have a real connection.
He showed seven knowledge tiers. At tier T7, models scored near zero percent, meaning they had no training data for those topics.

Gemini 3.1 Pro might actually be over 10 trillion, but since it hit a ceiling, the paper could not measure it directly.
This means that within a certain range, we can use training cost and data efficiency to estimate model size. But at the very top end, some models may be bigger than the numbers suggest.
In the original study, Bojie Li used about 1,400 real exam questions as the dataset. The idea was simple. By comparing how open-weight models scored on these questions, researchers could draw a curve and use it to guess the size of closed models.

One thing to note. The 90 percent prediction intervals in the original paper were extremely wide.
As many researchers pointed out, these wide ranges were only estimates. They should not be treated as facts.
undress ai remover
In short, we still do not know the exact size of GPT-5.5.

Bojie Li himself admitted that for the same task, results could span a 60x range. Any single point estimate would be dishonest.
ai nude generator
Still, IKP is a starting point, not the final answer.
The author honestly said he rushed an immature arXiv paper out the door, just to get the idea into the world.
The paper, code, dataset, and website were all built in four days, mostly using Claude Code, with no peer review before release. The flooring method and lambda value were chosen to maximize R2 on open-weight models.
We look forward to better work in the future.
Is the Scaling Law Dead
The collapse of this parameter myth is a wake-up call for the entire industry. The era of blindly worshipping big numbers is fading.
GPT-5.5 dropping from 10 trillion to 1.5 trillion does not mean it got weaker. It means OpenAI may have pulled off something far more impressive. Better data quality. Smarter parameter efficiency.
As Lawrence Chan put it in his summary, we still do not know exactly how many parameters GPT-5.5 has. But this method of probing knowledge capacity to reverse-engineer model size gives us a new path to peek inside the black box.
On the road to AGI, what we need may no longer be bigger hard drives. What we need is a smarter way to index what we already know.