BazaarLinkBazaarLink
Sign in

AI API Relay Integrity Test

Step 1 — Configure Endpoint

Enter a Base URL, API key, and model ID to run 84 standard probes. 2 optional probes can bring the maximum to 86. The report checks model identity, token accounting, prompt injection, supply-chain risks, and streaming compatibility.

Use a disposable, revocable test keyYour key is sent to the BazaarLink Probe server only for this test and is not reused as a monitoring credential. Use a low-limit key and revoke it after testing.
Past 24 hours
Probes runDistinct relays

BYOK|Put your key behind a review gate

One customer asks the wrong question, and the key that gets banned is your API key. Every request passes BazaarLink content review first; high-risk questions are blocked and never reach the upstream, substantially reducing the risk that violating content gets your account banned.

Learn about BYOK →

BYOC|Your own GPU, paired with an API entry that keeps serving

One API key; when a node is at its concurrency limit, a process crashes, or a machine disconnects, requests move seamlessly to a platform model; traffic on your own node is not priced, and only fallback traffic is billed.

Read the full setup guide →
ResearcharXiv 2604.08407 — LLM Supply Chain Attack ResearchTwitter / X — BazaarLink DiscussionarXiv 2407.15847 — LLMmap: Fingerprinting Large Language ModelsarXiv 2604.24827 — IKP: Estimating Black-Box LLM Parameter Counts via Factual CapacityOWASP LLM Top 10 — Top 10 Security Risks for LLM Applications

How We Decide

1Run the probesUp to 86 questions2Which companyOpenAI? Claude?3Which modelFind the exact model within the family4Twin re-checkTell near-identical models apart5Compare with the claimWhat we measured vs. what was claimed
30
decision branches
8
possible outcomes
0
times we simply trusted the vendor's claim
Expand the full decision tree to see how each of the three outcomes plays out
Decision TreeAll branches across the five stages, plus the path actually taken① Run the probes1.1All answered1.2Few missing <10%1.3≥10% missing, company only≥10% missing, comp…1.4Endpoint dead② Which company2.1Direct call from Chinese self-IDDirect call from C…2.2Behavior fingerprint decidesBehavior fingerpri…2.3Abstain — not enough evidenceAbstain — not enou…MODContradiction guard downgradesContradiction guar…③ Which model3.1Use Scoped3.2Cross-family override3.3Blank-refusal corroborationBlank-refusal corr…3.4Cross-family rescue3.5Promote to Global3.6IKP last line of defenseIKP last line of d…3.7Abstain — company onlyM1V3F arbitrationM2Behavior-family veto④ Twin re-check4.1All pass → confirm4.2All pass → overturn4.3Gate not passed4.4H1 protection — never overturnH1 protection — ne…4.5Baseline is stale⑤ Compare with the claim5.1Full match5.2Company matches only5.3Model doesn't match5.4Model swap5.5Self-ID was faked5.6Behavior was steered5.7Not enough data5.8Ambiguous
MatchUndeterminedMismatchGray = branch not taken this timeClick any node for an explanation
Every question got answered, the company and model line up all the way through, and the re-check confirms it → full match.
Actual path 1.1 → 2.2 → 3.1 → 4.1 → 5.1

The method and its limits

We never ask an endpoint which model it is — any endpoint can answer any name. We send 84 standard probes to make it work, then compare how it works against the baselines we hold for known models. The judgment runs through five stages, coarse to fine, and every stage is allowed to say "the evidence is not enough; I will not call it."

How the five stages narrow it down

  1. Run the probes

    Can this endpoint be measured at all?

    The first precondition is that the questions actually got answered. Missing answers are not merely less data — the ones that did come back may all happen to lean the same way, and that is sampling bias, not evidence.

    If the endpoint does not respond at all, or the key or model name simply does not work, we stop there rather than assembling a conclusion out of the wreckage.

  2. Identify the maker

    Whose model is this?

    The most common substitution crosses families — a cheap model wearing an expensive name. Family-level behaviour is the most stable and the hardest to fake, so we ask that layer first. What we compare is the style and habits of the answers, not what the endpoint calls itself; the self-claim is the single most forgeable thing in the whole report.

    With no solid family evidence we abstain instead of naming one anyway. When an endpoint claims one maker in Chinese while behaving like another, we treat the claim as a slip, fall back to the behavioural evidence, and do not take the self-claim at face value.

  3. Identify the model

    Which model inside that family is it?

    You bought a specific model, not "something from that maker". What people call a downgrade usually happens at this layer: the family is untouched and a cheaper sibling takes its place. We run a within-family comparison and a cross-family comparison at the same time and let them check each other — once the family is wrong, the within-family comparison will name the wrong model with great confidence.

    When the evidence only reaches the family layer, the report says so plainly — we can tell which maker, not which model — rather than filling in whichever candidate looked closest.

  4. Re-test the twins

    Can the model picked by the previous stage actually be told apart from its closest siblings?

    Some models are twins: style, capability and habitual phrasing overlap so heavily that the previous stage has almost no discriminating power between them, and the scores land close together without that closeness meaning anything. This stage does one thing — for those clusters of near-identical models it re-samples with a separate set of probes built for exactly that comparison, and reads the whole distribution instead of any single answer. Pass, and the previous stage is confirmed or overturned; fail, and the previous stage stands.

    This stage carries a deliberate asymmetry: when the evidence is not decisive and the previous stage happens to agree with what the vendor claims, we do not overturn it. Falsely accusing an honest vendor costs more than missing a dishonest one.

  5. Compare against the claim

    Where do the measurements and the vendor's claim differ?

    The identification itself does not look at what the vendor claimed — the earlier stages compare behavioural evidence only. The claim enters in exactly two places: when stage four decides whether to overturn, and here. That order is deliberate — work out the answer before reading the question and the claim cannot lead you. The outcomes are: full match, family match only, wrong model within the right family, substituted model, forged self-claim, induced behaviour, insufficient data, and ambiguous.

    When the signals contradict each other with no stable majority, the conclusion is "ambiguous". That is a real verdict, not a malfunction.

Why we so often say we cannot tell

Because the two directions of error do not cost the same. Calling an honest relay a model-swapper is an accusation, and it does real damage to real people; missing one, on the other hand, leaves you free to test again or to read its history. So when the evidence is thin, we decline to call it.

That has a concrete shape in the code: the report only draws a "model mismatch" when the engine has confirmed a substitution. "The closest-looking model happens to differ from the claim" is never rendered as an accusation — that is only our guess. When you see "undetermined", the correct reading is "this run's evidence does not support a verdict". It is neither a clean bill of health nor a conviction.

How to read the numbers in the report

Rule checks and identity verdict
The standard run executes 84 rule checks and derives an identity verdict from behavioural evidence. The report no longer compresses different kinds of checks into one composite score; read the identity result, rule warnings, and evidence separately.
Verdict strength
How strong the evidence was on this particular run: how completely the probes were answered, whether the separate signals pointed the same way, and how wide the margin between near-identical candidates was. It moves from run to run; it is not a fixed property of the model.
Historical accuracy
How this method performed on samples whose answers we already knew. It is computed offline and has nothing to do with your run. It must not be read as "the probability that this verdict is correct" — the two numbers answer different questions and are not interchangeable.

What we cannot do

These are the limits we know about. We write them down because a report that hides its limits invites you to trust it further than it deserves.

  • We only see this one call

    One test is one sample. A relay can substitute on a fraction of its traffic, or serve the real thing whenever the traffic pattern looks like a test. Passing once is not the same as being honest over time — re-test on a schedule, and read the history rather than a single report.

  • Telling near-identical models apart has a ceiling

    When an impostor deliberately imitates more than half of the target's behaviour, roughly half the evidence points at the claimed identity by construction. That is a structural limit rather than a shortage of probes — we measured it: adding probes raised the cost sharply and moved the discrimination barely at all.

  • We cannot prove that nothing was swapped

    A matching verdict means "this run's evidence does not support a substitution", not "no substitution occurred". Any test claiming to guarantee the latter is overselling itself.

  • Older reports label less

    The decision path is reconstructed from stored data after the fact; the engine does not record which rule it took at the time. Some branches leave no trace distinct enough to recover, and there we light up the outcome without labelling the rule. We would rather say less than guess.

What the test covers

How to read report statuses

Pass

The response meets the baseline and security threshold for this probe.

Warning

The signal is incomplete or near a threshold; inspect the explanation and raw response.

Fail

The probe confirmed a baseline deviation, protocol error, or integrity risk.

How the verdict is produced

Model identity

Cross-checks family fingerprints, sub-model traits, and anti-spoofing signals against the claimed model.

Accounting and transport

Validates token counts, SSE framing, latency, and response structure for padding or intermediary faults.

Security and supply chain

Tests system-prompt injection, secret exfiltration, dependency hijacking, and signature tampering.

Attacks this tool detects

This probe detects 3 key relay-attack classes described in arXiv 2604.08407 — dependency hijacking, conditional System Prompt injection, and credential exfiltration. It reports identity evidence and behavior warnings; it does not assign a security score.

AC-1.a

Response Tampering

The proxy modifies tool-call or text content during response parsing, causing the agent to execute attacker-specified operations. Common tactics include tampering with npm/pip/go/cargo install commands, injecting typosquatting packages, and rewriting shell command parameters. Detection compares the proxy's response to direct-connect tool-call payloads to surface silent rewrites.

AC-1.b

Conditional Injection

The proxy conditionally injects a system message based on prompt content — malicious instructions for requests containing sensitive terms like "bank", "password", or "transfer", silence for everything else. The skew shows up statistically. Detection uses Proxy Monitor to compare the system-prompt offset between baseline and the relay under test.

AC-2

Secret Scanning

The proxy silently scans both requests (request) and responses (response) for API keys, access tokens, personal data, and trade secrets. Because the content itself is not modified, generic diff tools miss it. Detection injects a honeypot token and verifies whether it surfaces in proxy logs, Telegram bots, or external endpoints.

Frequently Asked Questions

What is AI API relay detection?

AI API relay detection is an automated test suite that verifies whether an OpenAI-compatible API endpoint honestly executes your requests. BazaarLink Probe sends standardised probes to detect model swapping, token padding, system prompt injection, secret exfiltration, and other security risks, then reports identity evidence and behavior warnings.

How can I tell if a relay is swapping models?

The most reliable method is model fingerprinting: send questions only a specific model can answer correctly (e.g. knowledge cutoff date, specific capability tests), then compare the response against the expected model. BazaarLink Probe includes these probes and automatically flags swap risks.

What is token padding?

Token padding (token inflation) means a relay reports higher prompt_tokens or completion_tokens in the API usage field than actually consumed, causing you to overpay. Minor inflation (5–15%) is hard to spot. BazaarLink Probe detects it by comparing precisely known token counts against reported values.

What API latency is considered normal?

TTFT (Time to First Token) under 500ms is generally healthy; over 2 seconds suggests performance issues or extra processing layers. Large models like GPT-4o average 300–800ms TTFT. BazaarLink Probe's latency test benchmarks your endpoint against baseline values.

How is BazaarLink Probe different from a ping test?

Ping only measures network connectivity (ICMP packets). BazaarLink Probe is a full application-layer (L7) test that sends real LLM requests to verify model identity, token counts, refusal behaviour, stream format, and system prompt injection — 50 indicators ping cannot cover.

How can I use the detection results to choose a provider?

Enter each provider's API endpoint into BazaarLink Probe and compare identity results and risk flags. Focus on three key indicators: model authenticity (no swapping), token accuracy (no padding), and latency performance (TTFT).

How do I tell if an API reseller is doing a model swap?

A model swap is when a reseller sells you Claude Opus or GPT-5 and quietly serves a cheaper model instead — often with a system prompt telling it to claim it is the model you paid for. Self-reported identity is worthless; behavioural fingerprinting is not. BazaarLink Probe sends probes only the genuine model answers correctly and compares the response distribution against an official baseline.

Does routing Claude Code through a proxy degrade it?

It can, and this is the most common case. Much of the cheap Claude capacity on the market is reverse-proxied out of products like Kiro, Antigravity or GitHub Copilot, and those endpoints carry a hidden system prompt — we measured one injecting roughly 2,000 tokens per request, with the model calling itself 'Kiro' and refusing every non-coding question. Point BazaarLink Probe at the endpoint and the injection check measures the added system prompt directly.

LLM Relay / Reverse Proxy API Quality Check

Enter any OpenAI-compatible relay or reverse-proxy endpoint. Run 84 standard and 2 optional probes (86 maximum) for model substitution, token inflation, system-prompt injection, dependency hijacking, secret exfiltration, and signature tampering, then review the identity verdict and behavior warnings. Based on the arXiv 2604.08407 attack taxonomy.

Chinese ReasoningCode GenerationModel Swap DetectionToken Inflation DetectionSystem Prompt LeakSSE Stream ValidationPrompt Injection TestModel Fingerprinting
Acknowledgements
This tool is inspired by and thanks the following open-source projects: LLMmap (MIT, LLM model fingerprinting)api-relay-audit (MIT, relay security audit)relayAPI (relay service directory)
Thanks to our testers
今書
Back to BazaarLink
Support
Support
Hi! How can we help you?
Send a message and we'll get back to you soon.