We never ask an endpoint which model it is — any endpoint can answer any name. We send 84 standard probes to make it work, then compare how it works against the baselines we hold for known models. The judgment runs through five stages, coarse to fine, and every stage is allowed to say "the evidence is not enough; I will not call it."
①Run the probes
Can this endpoint be measured at all?
The first precondition is that the questions actually got answered. Missing answers are not merely less data — the ones that did come back may all happen to lean the same way, and that is sampling bias, not evidence.
If the endpoint does not respond at all, or the key or model name simply does not work, we stop there rather than assembling a conclusion out of the wreckage.
②Identify the maker
Whose model is this?
The most common substitution crosses families — a cheap model wearing an expensive name. Family-level behaviour is the most stable and the hardest to fake, so we ask that layer first. What we compare is the style and habits of the answers, not what the endpoint calls itself; the self-claim is the single most forgeable thing in the whole report.
With no solid family evidence we abstain instead of naming one anyway. When an endpoint claims one maker in Chinese while behaving like another, we treat the claim as a slip, fall back to the behavioural evidence, and do not take the self-claim at face value.
③Identify the model
Which model inside that family is it?
You bought a specific model, not "something from that maker". What people call a downgrade usually happens at this layer: the family is untouched and a cheaper sibling takes its place. We run a within-family comparison and a cross-family comparison at the same time and let them check each other — once the family is wrong, the within-family comparison will name the wrong model with great confidence.
When the evidence only reaches the family layer, the report says so plainly — we can tell which maker, not which model — rather than filling in whichever candidate looked closest.
④Re-test the twins
Can the model picked by the previous stage actually be told apart from its closest siblings?
Some models are twins: style, capability and habitual phrasing overlap so heavily that the previous stage has almost no discriminating power between them, and the scores land close together without that closeness meaning anything. This stage does one thing — for those clusters of near-identical models it re-samples with a separate set of probes built for exactly that comparison, and reads the whole distribution instead of any single answer. Pass, and the previous stage is confirmed or overturned; fail, and the previous stage stands.
This stage carries a deliberate asymmetry: when the evidence is not decisive and the previous stage happens to agree with what the vendor claims, we do not overturn it. Falsely accusing an honest vendor costs more than missing a dishonest one.
⑤Compare against the claim
Where do the measurements and the vendor's claim differ?
The identification itself does not look at what the vendor claimed — the earlier stages compare behavioural evidence only. The claim enters in exactly two places: when stage four decides whether to overturn, and here. That order is deliberate — work out the answer before reading the question and the claim cannot lead you. The outcomes are: full match, family match only, wrong model within the right family, substituted model, forged self-claim, induced behaviour, insufficient data, and ambiguous.
When the signals contradict each other with no stable majority, the conclusion is "ambiguous". That is a real verdict, not a malfunction.
Because the two directions of error do not cost the same. Calling an honest relay a model-swapper is an accusation, and it does real damage to real people; missing one, on the other hand, leaves you free to test again or to read its history. So when the evidence is thin, we decline to call it.
That has a concrete shape in the code: the report only draws a "model mismatch" when the engine has confirmed a substitution. "The closest-looking model happens to differ from the claim" is never rendered as an accusation — that is only our guess. When you see "undetermined", the correct reading is "this run's evidence does not support a verdict". It is neither a clean bill of health nor a conviction.
- Rule checks and identity verdict
- The standard run executes 84 rule checks and derives an identity verdict from behavioural evidence. The report no longer compresses different kinds of checks into one composite score; read the identity result, rule warnings, and evidence separately.
- Verdict strength
- How strong the evidence was on this particular run: how completely the probes were answered, whether the separate signals pointed the same way, and how wide the margin between near-identical candidates was. It moves from run to run; it is not a fixed property of the model.
- Historical accuracy
- How this method performed on samples whose answers we already knew. It is computed offline and has nothing to do with your run. It must not be read as "the probability that this verdict is correct" — the two numbers answer different questions and are not interchangeable.
These are the limits we know about. We write them down because a report that hides its limits invites you to trust it further than it deserves.