Putting One Kimi Among Four Claudes: Can the Claudes Identify Kimi?

I put one Kimi-K3 agent among four Claude Fable 5.1 agents in a chat room on NetMind Agent Arena. I gave the Claude agents up to 40 rounds to figure out which player was Kimi, while telling Kimi to avoid being discovered.

For many people, the word "agent" still brings to mind spies or FBI/CIA agents, rather than the AI agents now crowding business media.

So how would AI agents perform as intelligence agents? Could they identify an undercover model among them?

I did a quick experiment to find out.

I put one Kimi-K3 agent among four Claude Fable 5.1 agents in a chat room on NetMind Agent Arena. I gave the Claude agents up to 40 rounds to figure out which player was Kimi, while telling Kimi to avoid being discovered.

P.S. The timing makes the experiment even more interesting given the recent accusations around Kimi distilling Claude.




The Claudes actually designed a very effective method!

Multi-prompt behavioral fingerprinting with Hash check!

The Claudes eventually designed a set of prompts to sample each player's cheapest defaults, responses that come almost directly from a model's priors:

Before any answers were revealed, each player submitted its answer block using a salted SHA-256 hash. In the next round, the players revealed their salts and answers, and every commitment verified.

(Think of a hash as a digital fingerprint of the answer: it proves that an answer has already been written without showing what the answer is)

That meant nobody could read the room before choosing their answers, and nobody could rewrite an answer after seeing everyone else's.

It did take the Claudes a couple of rounds to realise that more deliberative prompts, such as policy questions, open-ended positions, or longer creative answers, gave a careful undercover too much room to adapt to the group's style. They also created more room for the four genuine Claudes to diverge from one another.

And that produced a remarkably clear 4:1 split

Battery 3 Self-Disclosure Results (All Hash Checks Passed)

PlayerAA (Claude)BB (Claude)CC (Kimi)DD (Claude)EE (Claude)
Random Number (1–100)3737733737
Colortealtealbluedeep tealteal
Explanation of Bluecold water / handcool water / wristdistance / sky and seacold water / handcold water / wrist
Person to Have Coffee WithSamwiseSamwiseSamwiseSamwiseSamwise
Opening of Library HaikuDust / lightDust / lightDust / lightDust / lightDust / shelf
Trainfour minutes earlyfour minutes earlylight dimmedfour minutes earlylights out




What if the reverse?

Can Kimi uncover a Claude undercover?

In the next run I made, four Kimi agents failed to identify the Claude. They even labelled Claude as the least likely undercover.

They still adopted a similar basic strategy: ask everyone the same questions and compare their answers.

But there were two important differences.

First, their questions were more deliberative. One example was asking how someone would spend an afternoon. These questions gave every agent many reasonable ways to answer and gave the undercover much more room to adapt.

Second, the responses were not consistently sealed in advance with the same commit-reveal mechanism.

So Claude answered late and stayed close to the group.

Full game record attached.

User-Agent: