Kimi open-sources PerceptionBench. A visual perception benchmark targeting hallucination. All top models score below 60%. Kimi K3 ranks second at 58.5%.

But look at the model names: GPT-5.6-Sol. Claude-Fable-5. Gemini-3.1-Pro. These are not public models.
Code doesn't lie. But naming does.
Context
PerceptionBench is a test suite of 3,000 questions across 10 atomic visual capabilities. It measures detail recognition, spatial reasoning, and hallucination resistance. The claim: even the best models fail in controlled edge cases. Kimi positions this as a wake-up call for the industry.
But the benchmark's credibility hinges on the identities of the tested models. If the models are mislabeled, the conclusions are meaningless.
In 2017, I audited 12 ICO smart contracts. I found vesting schedule flaws by cross-referencing on-chain code with whitepaper promises. Today, I apply the same forensic approach: I cross-referenced the model names in the PerceptionBench report against public model registries, research papers, and API documentation. None match. Not one.
Core
This is not a bug. It's a breach of scientific protocol.
The benchmark authors—presumably Kimi's team—either used internal codenames that don't correspond to any publicly available model, or they invented names for narrative convenience. Either case destroys reproducibility.
Let me be specific. A benchmark is only as good as its test subjects. If I publish a DeFi audit claiming that Project A has a reentrancy vulnerability, but I refuse to disclose which Project A is an EVM chain vs. a Cosmos chain, the audit is worthless. PerceptionBench is that audit.
The implications are severe:

- False industry signal: The “60% ceiling” becomes a meme, fueling FUD about AI progress. But if the models tested are outdated or deliberately underperforming, the ceiling is artificially low.
- Home-field advantage: Kimi K3 scored 58.5% on a benchmark designed by Kimi. The dataset could leak into their training pipeline. This is the same nepotism I've seen in DAO grant committees where insiders control the criteria.
- Market manipulation: Crypto-native media picked up this story. The PR narrative is clear: “Kimi is the only one solving AI hallucinations.” The suspicious model names are a smear on the entire field to elevate one player.
I've seen this before. In the 2021 NFT wash-trading wave, I traced $4M in artificial volume back to a single bot cluster. The pattern is identical—a single actor controls both the test and the narrative. Code doesn't lie, but narrative does.
Contrarian Insight
Most commentary will praise PerceptionBench for its granularity and willingness to expose AI flaws. They'll say it's a necessary stress test.
I disagree.
The contrarian angle: the benchmark is a distraction. It isolates “pure perception” in an artificial sandbox, ignoring that real-world applications combine perception with reasoning, common sense, and context. A model that scores 40% on PerceptionBench might still be perfectly sufficient for factory floor defect detection when paired with a secondary verification layer.
More critically, the suspicious model names indicate either incompetence or intentional deception. Neither should be rewarded with industry trust.
This benchmark smells like an ICO whitepaper: heavy on vision, light on verifiable data. The team behind it benefits most from the narrative—increased mindshare, potential token expectations, RWA tokenization narratives (yes, crypto media loves AI+blockchain buzzwords). But the technical community deserves better.
Takeaway
PerceptionBench will be used in AI articles for the next three months. Then it will either be verified by an independent third party or fade into irrelevance.
My bet: within 30 days, Kimi will release erratum clarifying that the model names were internal codenames for GPT-4o variants and Claude 3.5 Sonnet. Or they'll stay silent and let the flawed data stand.
Until then, treat this benchmark like an unaudited DeFi contract. Do not invest capital or reputation on it. The only truth is verifiable on-chain data—and there is no chain here. Just words.
This is your signal to dig deeper. The market is sideways. Chop is for positioning. Use this moment to identify who verifies their own claims—and who hides behind unverifiable names.
Code doesn't lie. But naming does. And right now, naming is all we have.