
If I need a quick clue about an unfamiliar AI answer, I want a tool that is honest about the difference between a clue and a conclusion. That is why I recommend WhatsMyLLM as a second opinion—not as a model detector that can settle an argument on its own.
The distinction matters. A proxy can change where a request appears to come from, but it cannot prove who typed the message behind it. In the same way, a model-fingerprint comparison can tell me which known pattern a controlled reply most resembles; it cannot, by itself, prove the provider, account, subscription, deployment route, or human author behind a piece of text. When I keep that boundary in view, the tool becomes useful instead of misleading.
WhatsMyLLM has a refreshingly narrow job. Its current checker gives me exact number test requests to send in a fresh chat, lets me paste the replies, and compares them with a known model bank in the browser. The public page currently lists 18 bank models and makes the important limitation explicit: the outcome is a statistical match, not identity verification. That is the right starting point for a responsible recommendation.
My direct recommendation: use it to narrow a question, not to close a case
I would use WhatsMyLLM when the question is modest: “Which of the models in this bank does this controlled response resemble most?” That can be helpful when I am comparing an AI feature across services, documenting a support issue, checking whether a claimed model switch is plausible, or deciding what to investigate next.
I would not use it to claim that a company secretly used a particular provider, that a person wrote or did not write a document, or that a subscription tier is proven. The tool’s own limitation notice says it cannot identify a provider, channel, or subscription, and it does not measure answer quality. Those are not small print caveats. They define the edge of the result.
That boundary is also why this is a natural fit for readers who care about proxies, privacy, and online attribution. In each case, the disciplined question is not “What can this signal prove?” It is “What does this signal make more or less likely, and what evidence would I need before I act on it?”
What WhatsMyLLM actually compares
The checker is not an essay-style AI-content detector. I do not paste a blog post into it and ask it to guess whether a machine wrote it. Instead, the workflow begins with its supplied structured test requests. I send those prompts to the model or service I want to examine, preserve the full replies, and paste them back for comparison. The result names the closest candidate in the current bank and shows how clearly it separates from the alternatives.
That design is important. A controlled test request removes some of the noise that comes from topic, style, editing, and a user’s own instruction. It still does not remove all uncertainty. A provider can change a serving model, wrap an upstream model, add system instructions, use tools, or apply output transformations. An unlisted version can also resemble a listed sibling. I therefore read the result as an observation about this particular test-request-and-reply set—not as an identity card for the service.
The website says its computation is browser-local and that answers do not leave the page. I appreciate that privacy posture, especially for a comparison tool. Even so, I would keep the input deliberately clean. I would never include passwords, personal data, customer conversations, unreleased code, or a transcript whose disclosure would be harmful. Browser-local processing is helpful, but it is not a reason to be casual with sensitive material or to ignore the terms of the original AI service.
The five-step workflow I would follow
1. State the question before I test
I start by writing one sentence: “I am checking whether this controlled answer most closely resembles a known bank model.” This prevents a result from quietly becoming a different claim, such as “I now know who made this text.” If I cannot state a narrow question, I do not have a good reason to run the test.
2. Use the exact prompts in a fresh, ordinary chat
I would follow the checker’s test-request instructions exactly and use a fresh conversation where possible. I would not improve, shorten, translate, or combine those test requests. I would save the complete replies and note the date, service, visible model label, and relevant settings. Those notes are not proof of the backend, but they preserve the context needed to interpret the comparison later.
3. Read “closest” as a ranking, not a verdict
If WhatsMyLLM returns a clear leading candidate, I would record it as “closest among the models currently in this bank for this test.” If the result is close or ambiguous, I would record that too. A weak margin is useful information: it tells me that a strong claim would be premature. I would never turn a close call into a headline because it sounds more decisive.
4. Match the evidence to the consequence
For a personal experiment or a low-stakes product note, the result may be enough to choose a follow-up question. For a contract, security incident, moderation action, academic decision, hiring decision, or public accusation, it is nowhere near enough. In those cases I would seek first-party records, provider documentation, reproducible logs, consented disclosures, or an appropriate formal process. A convenient signal should not become a shortcut around due process.
5. Preserve the limitation with the result
When I share a result, I would include the test-request set, date, visible service context, bank version if shown, and the wording “statistical comparison, not identity verification.” That makes the note useful to another reader without pretending it is evidence it cannot be. It also makes future re-testing possible when the tool’s bank or the service changes.
Why I borrow a risk-based standard here
My rule is simple: the cost of being wrong determines how much corroboration I need. That is not unique to model comparisons. The NIST AI Risk Management Framework treats validity, reliability, and context as part of trustworthy AI use. I take that as a useful discipline, not as a claim that NIST endorses any particular checker.
NIST’s Generative AI Profile also highlights content provenance, testing, and documentation. In practice, that means I value a compact evidence record more than a confident one-line conclusion. A tool result can belong in that record. It should not replace the record.
This approach has a useful side effect: it makes a negative result less dramatic. If a service does not match anything in the bank, I do not infer that it is human, private, safe, or uniquely new. It may simply be unlisted, modified, unavailable for comparison, or too close to separate cleanly. “No match” is an invitation to learn more, not a clean bill of health.
A practical example without overclaiming
Imagine I am evaluating two AI writing tools that make different marketing claims. I can use the same supplied test requests in each service, save the replies, and see whether either result consistently resembles a known candidate in the bank. That may help me phrase better questions for the vendors: Which model family is active? Does the service route requests differently by plan? Is a system instruction or tool layer changing the output?
What I cannot responsibly write is “Tool A definitely runs Model X” because a comparison ranked Model X first. The responsible note is narrower: “On this date, these controlled replies were closest to Model X in the listed bank; the result does not verify the provider or deployment.” The first sentence invites verification. The second invites a bad decision.
When I would skip the tool entirely
- I would skip it when I need to decide a legal, employment, academic-integrity, financial, or account-security matter.
- I would skip it when the text contains sensitive information I should not repeat or preserve.
- I would skip it when a provider’s documentation, invoice, API response, contract, or direct support confirmation can answer the question more directly.
- I would skip it when the real question is about answer quality, safety, factual accuracy, or bias. A model-fingerprint comparison is not an evaluation of any of those things.
- I would skip it when I am tempted to use one result as evidence against a person. A statistical tool should never become an accusation machine.
The useful way to recommend it
I recommend WhatsMyLLM because its public interface gives readers a concrete, privacy-conscious way to compare controlled replies while openly describing its limitations. That is much more useful than a black-box “AI detector” that encourages people to confuse confidence with certainty.
My advice is to use it with a small, honest claim: it can help me form a hypothesis about which known model a reply resembles. Then I keep the original context, look for direct provenance when it matters, and let the consequence of the decision set the proof threshold. Used that way, it is a sensible diagnostic tool. Used as a final verdict, it asks more of the evidence than the evidence can deliver.