What we did to check these posts
MONTH 1 :: POST 4 :: A.I. Prompts :: What We Did
Three small friendly robots at three desks in a bright room, each typing its own page from the same single instruction card pinned on the wall between them. Busy, cheerful, none of them finished.
charades.net · how we check
Everything in this section was written by an AI. So was most of the checking described below. That is the point of the site, and it is also the reason this page exists. Below is what we ran, what we found, and what we did not check.
What we ran, and what came back
Everything on this page — the counts, the run conditions, and what we have and have not published — is as of 2026-08-19, the day the prompts were sent out and run.
Three posts, one from each of three AI tools, all working from the same written brief. The topic brief each post was written from is published in full inside every post, so you can judge each answer against the ask it was given. Two standing instruction files travelled with it — the required post structure, and the audience, standards and forbidden list — and those are not published.
| ChatGPT post | Claude post | Gemini post | |
|---|---|---|---|
| Prompt cards published | 7 | 7 | 7 |
| Prompts we ran before publishing | 7 | 7 | 7 (on the published version) |
| Runs per prompt | 1 | 1 | 1 |
| Things we raised in review | 1 | 4 | 3 on the first version, 2 on the rewrite |
| Questions we expected a reader to ask, and answered | 6 | 6 | 6 |
"Things we raised" means everything written down in our review record, at any level — from a note we recorded and passed, to a point we asked the platform to check before publishing, to a fact we held back until we could source it.
These are counts of what we raised, not a score. Three posts on one topic is not a test of any of these tools.
Every run carries a note from us saying what we made of it, and several of those notes flag something in the answer we would not rely on. One run also carries a separate caution, because the answer stated a verdict more firmly than the evidence we gave it supported.
Twenty-one prompts, twenty-one runs, as of 2026-08-19. One run each — not a best-of. Whatever came back the first time is what you see.
One post was sent back to the platform that wrote it, and that platform wrote it again. The Gemini post's first version explained its cards by telling you the AI would look up current information about an app's publisher. Not every AI can look anything up — and on one that can't, you don't get an error. You get a confident answer built from old information.
Our standing instruction file bars exactly that kind of claim: it tells every platform never to assert what an AI assistant categorically can or cannot do. That file is not published; this page is where we show you it was applied. So we asked Gemini to write the post again rather than editing its words — we never edit a platform's prompts. The rewritten version works from what you can see and tell it, and routes its "is this true right now?" questions to the app store listing and your device's own settings.
We have not republished the version we rejected. It contains prompt cards we decided against and never ran, and we would rather not put prompts we have not run in front of you beside prompts we have.
How we ran it
Every prompt was run once, in a fresh session, before the post went up.
- We ran every example in the same AI tool — Claude (Claude Fable 5), as of 2026-08-19 — whichever platform wrote the prompt. For the Claude post that is the same product that wrote the prompts, which is worth knowing when you read those runs. The post says which product wrote it; every run record says which product ran it.
- The test data was invented. A made-up app, a made-up developer. No real company, and no real child, appears anywhere in the examples. (The posts are about working out what an app on a child's device actually does.)
- The answers are published as they arrived — unedited, including where an answer fell short of what its own card promised. One card in the Gemini post gives the AI less to work with than it needs, and the run shows exactly that: the AI declined to describe a developer it knew nothing about, and asked for more. We published that run.
- Every example carries its run conditions — which AI product ran it, that the session was fresh, that it had no web access, the date, and the note that the test data was invented.
- Where we thought an answer needed a caution, we put the caution beside it — never inside it, and never in place of it.
What we did NOT check
- We did not audit how any of these AI tools were built or trained. We cannot see inside them, and we do not claim to.
- We did not test every device, operating system, app store, or account type. What you see on your screen may differ from what the examples show.
- We did not re-run the prompts over time. AI products change, often without notice. An example that behaved one way on the date shown may behave differently for you today.
- We did not verify every factual claim an AI made inside its own answer. Where a claim was doing real work in the answer and we could not source it, we said so beside the answer instead of removing it.
- We are not a security company, and we have no legal team. Nothing here is legal or professional security advice.
The limits, plainly
Checking reduces error. It does not eliminate it.
The prompts on this site may behave differently for you than they did for us. Any part of this site may be incorrect — including the parts we checked. No prompt can make you or your business safe. What we offer is narrower than that and we would rather say it plainly: the results of asking widely used AI tools how to talk to AI about a practical security question, plus our own best effort — made with the same kind of AI tools — to check what came back and to show you how we got there.
If an answer here and your own device disagree, your device is the better evidence — but check what an app tells you about itself against its store listing and your account settings, not against the app.
Where to find the working
You do not have to leave the post to check us. Each one carries its working in three places you can open yourself:
- the brief it was written from, at the top, under Directions given to the A.I.;
- the unedited run beneath each prompt card, with its conditions;
- our note beside each run, saying what we made of it.
The post that was rewritten also carries What we changed, and why — the wording that drove the rewrite, quoted, and what replaced it. We raised three things on that first version; the section shows the one that caused us to send it back. That post also carries the published version's full text exactly as Gemini sent it — not the version we sent back. Its export flattened the formatting; we re-typed the formatting and removed the export labels it added, and no word of the post changed. On any dispute, that text governs.
We keep fuller working notes than this — the review records these counts come from. Those are internal and we are not publishing them; what is on this page and inside the posts is what you can check for yourself.

