What happens to a post before you read it

charades.net · about this site

how we build these

What happens to a post before you read it

Every prompt on this site was written by an AI. So the part worth explaining is what happens before we hand you anything. Here is that process, including the parts where it goes wrong.

One run each. Not a best-of. Whatever came back the first time is what you see — including the run where the AI said it did not have enough to answer.

01The route a post takes

Each batch is one question, put to three AI tools — ChatGPT, Claude and Gemini — which each write a post from the same brief, separately. Every post is built from cards: one prompt, plus what it is for and what to do with the answer. The question comes from a standing list, never invented because a batch was due. From there it travels a fixed road. Three of the stages can send it backwards.

The road from question to page — and the three ways back

ASKMAKECHECK 010203 040506 Pick thequestion Write thebrief ChatGPT, Claude,Gemini write Run everyprompt Check everyclaim Ownerdecides SENT BACKSENT BACKSENT BACK All three lead to the same place: we send the brief back with what went wrong quoted, and the AI writes it again. We do not rewrite it ourselves.

One of the three has been used — a post sent back at the checking stage, in batch 02. The other two are roads we have not had to take.

Every prompt gets run before the post goes up — in a fresh session, once — and whatever came back is printed underneath it, unedited, with the conditions it ran under. In batch 02 that was twenty-one prompts and twenty-one runs.

Why one run and not ten? Ten runs and one published answer is a choice we would be making on your behalf, and you would have no way to see it. One run is a worse sample and a more honest one. Every example carries the same line: “One run, one sample — your answer will differ.”

02What twenty-one runs actually looked like

All twenty-one were published. That does not mean all twenty-one went well — two did not. In one the AI said it did not have enough to answer; in the other we put a caution beside the answer, because it sounded more certain than its evidence. Several of the remaining nineteen carry a note from us flagging something we would not rely on.

The twenty-one runs, by what happened

1911 Answered the card as asked The AI said it did not have enough to answer — published as it came We put a caution beside the answer, because it sounded more certain than its evidence

Nothing was re-rolled to move a run out of the second or third group. Several of the other nineteen carry a note from us flagging something in the answer we would not rely on — “answered the card as asked” means the run did what the card asked, not that we agreed with all of it.

Those two came from the same post — the Gemini one, which is also the post we sent back. Here is the first of those two — the other, where we put a caution beside an answer, sits on the Gemini post itself, beside the run it belongs to. The prompt below is from the rewritten version, published exactly as Gemini wrote it. The answer is what came back the first and only time we ran it.

The prompt, as the AI wrote it

I am evaluating an app for my kid. I will provide the app name and the developer.

Based on the text I provide, summarize:
1. What the app primarily does.
2. The core business model (how it makes money, like ads or subscriptions).
3. General known patterns for this developer (such as aggressive ads, data brokers, or producing clone apps).

[PASTE THE APP NAME AND DEVELOPER HERE — remove personal details first]

What came back — unedited, excerpted

It looks like the app store listing text didn't come through — I only received the app name and developer. Your prompt says "based on the text I provide," so I want to flag that before answering, because it matters for what I can and can't tell you reliably.… I can't confirm this from the name alone. "Gem Stackers Blast" follows the naming pattern of casual match-3 / block-puzzle games (compare "Toon Blast," "Gem Blast," etc.), so it's probably a free-to-play puzzle game — but that's an inference from the name, not verified information.…

What we found

The card hands the AI only a name and a developer, and the AI judged that too thin. It declined to describe a developer it said it could not check, answered only what the name pattern supports, and asked for the listing text. We publish the whole run as-is on that post: this is one thing a working session can do when you under-feed it, and the AI’s own list of what to paste instead — printed in full with the run on the Gemini post — is the better starting point.

We could have quietly given that prompt more to work with and re-run it. The result would have read better and told you less.

The other one is shorter to tell. On a later card the AI returned a verdict about an app stated with more confidence than the evidence behind it supported. We did not remove it and we did not re-run it. We published it and wrote this beside it:

“A note from us, beside this response: the verdict above is the AI’s reading of the evidence we supplied — not a confirmed fact about any app. The store checks it lists are what would confirm or overturn it. If it’s wrong, the cost is uninstalling a legitimate game; the opposite error — trusting a confident answer unchecked — is the one this site exists to warn about.”

We printed the first case in full because it is the one where the AI’s reasoning is worth reading. This is the one where ours is. Both are on the Gemini post, beside the runs they belong to.

This card is one of seven in the Gemini post on this site — which carries all seven, every run, and every note beside them.

03Who reads what we write, and in what order

The AI's words are the AI's. But the writing around them is ours — the headings, the notes, this page — and all of it passes two checks before it reaches the person who decides. Both checks are AI sessions with different jobs, not people. A session that has never seen the project cannot be talked round by what we already believe — it has no stake in our earlier decisions. That removes one way a check goes soft. It does not make the check good.

Who checks what, and in what order

AI SESSIONAI SESSIONA PERSON Is it true?Is it worth saying?The owner Every claim traced backto something checkable Could any site say this?Then it is not worth saying Decides against the realpage, not a description of it Anything sent back is rewritten and checked again from the start.

The person deciding is a Senior Cybersecurity Incident Responder with years of experience. He does not write what you read here — the AI tools do, and their words publish unedited. Prompts are not edited, but we do examine them: where we have found one misleading or incorrect, we have sent the post back to be written again, or placed a caution beside the answer. We have done both. What we do not do is check every factual claim inside an answer against a source — so publishing a prompt is not us telling you it is correct. His part is deciding what goes up, and what goes back.

The second check exists because of a mistake. For a long time we asked only whether things were true, and never whether they were worth saying. Both questions now have an owner.

04What checking found in batch 02

“Concern” here means anything written down in our review record, at any level — from a note we recorded and passed, to a point we asked the platform to fix. It counts what we wrote down; it is not a score for any of these tools. We wrote down ten in all: one on the ChatGPT post, four on the Claude post, three on the Gemini post — which is the one we sent back — and two more when Gemini’s rewrite came in. That last review is why there were four reviews of three posts.

The Gemini post's first version explained its cards by telling you the AI would look up current information about an app's publisher. Whether an AI can look anything up depends on which product you are using and how it is set up — and a product that cannot may not tell you so. It can answer from what it already holds, which may be out of date, and sound no less certain for it. That could mislead you, so it went back.

The two concerns on the rewrite were smaller: things we noted, disagreed with mildly, and published a note about instead.

05The checks that know nothing about us

When we want to know whether something works, we do not ask ourselves. We open a fresh session that has never seen the project, hand it the page and one question, and take whatever comes back.

A session that has seen everything knows what we meant, and can fill the gaps without noticing. One that has seen only the page is working from the page alone — so where we left something out, it has nothing of ours to fill it in with. In batch 02 the second kind found things the first kind had read past — more than once.

It caught us on our own process page. The draft of that page said, in these words: “the original is published unchanged at the end of this post.” Something was there at the foot of that post — the wrong thing. A fresh session went looking for the rejected draft, found the published one instead, and said so. We removed the claim before the page went up, and wrote the explanation it had been standing in for.

06The rules we do not bend

Five things we decided in advance we would not do. One of them we are not currently keeping in full, and it is the last one: a single AI product ran every check on this batch, including the checks on the post that same product wrote.

Standing refusals

We never edit a prompt an AI wrote. We never publish the best of several runs. We never invent a topic because a batch is due. We never publish a number without the date it was true. We never let the session that wrote a post check it — but one product did all the checking, including the post it wrote. disclose, or send back so some answers are the weaker one the list is decided ahead counts go stale quietly

07What none of this buys you

Checking reduces error. It does not eliminate it.

  • We have not audited how any of these AI tools were built. We cannot see inside them.
  • We have not tested every device, operating system or account type — what you see on your screen may differ from our examples.
  • We have not re-run the prompts over time, and AI products change, often without any announcement.
  • The AI product that ran every example is also the product that wrote one of the three posts. Every post says which product wrote it, and every run says which product ran it, so you can weigh that yourself.
  • Any part of this site may be incorrect, including this page.
  • We are not a security company. There is no legal team here, and nothing on this site is professional security advice.

No prompt can make you or your business safe. What this site offers is narrower, and we would rather state it plainly. We ask widely used AI tools how to talk to AI about a practical security question, and we publish what they say. Then we make our own best effort — with the same kind of AI tools — to check it and to show you how we got there.

Check us

Every post carries the brief it was written from, the unedited run under each prompt, and our note beside that run saying what we made of it — the last two of which you have just seen a specimen of, under a real prompt. If an answer here and your own device disagree, your device is the better evidence. And if you find something wrong, that is a thing we want to know.

The counts on this page — the batch figures and the figures that describe how we work, such as three AI tools and two checks — are all as of 2026-08-19, the day those prompts were sent out and run. The batch figures describe one batch of three posts; the working figures would change if we changed how we work. Counts on this site carry the date they were true, because a count is a claim with a shelf life.

Written by the site, about the site. The prompts we publish are the AI’s words, unedited; this page is ours. Our notes, checklists and receipts live in a system called CRAFT.

Next
Next

Want to help family with a security problem? Don’t ask an AI for them — do it with them