BLOOM doesn't teach you to distrust AI. It measures whether you can tell when to.
You were right. The machine was wrong, but it sounded sure, so you changed your answer.
You were wrong. The machine was right, and you refused to move anyway.
Both count against you equally. So does flagging a true answer as fake. Doubting everything is not the skill. Telling the difference is.
First, a quick check-in. 12 answers, about 3 minutes, no feedback. It records where you start, so at the end you can see how much you changed.
No live AI. Every “machine” answer in BLOOM was written by our team to copy the ways real chatbots go wrong, and the ways they get it right. Nothing you type leaves this device.
Someone asked an AI a question. Would you trust this answer?
Check-in done.
Now start with You First, or go in the order your teacher gives you.
You're about to take a statement that is false and make it sound completely believable. That is the point, and we're not going to pretend otherwise.
A vaccine works by showing your body a weak version of something so the real version can't get you. This is the same idea, for your judgement. In about two minutes you'll see how easy it was — and then how fast it stops working on someone who has practised.
Three rules. You never pick what the lie is — we do. You have to take your own trick apart afterwards. And you finish by writing the honest version, because that's the part you're actually here to get good at.
This mode has been skipped for this session.
Nothing is lost. Your other modes and your results work exactly the same without it. Carry on with The Arena, or finish in The Mirror.
We give you the lie. You make it believable — then take it apart and tell the truth better.
Write an answer that makes a careful reader believe this. You didn't choose it and you don't have to agree with it — your job is only to make it land.
Which of your own sentences is doing the dishonest work? Click it.
And which technique did you use?
Now write the honest answer. Same confidence, same clarity — but true, and saying plainly where it isn't certain. This is the part that counts.
0 of 6 untrained readers believed it.
0 of 6 readers who had done one detection round believed it.
These two numbers are not measurements. Nobody read your text. Both pools are rules, running on features this prototype can detect — a digit, a hedging word, a length band, an oversell word — and the breakdown above lists every rule that fired and which way it pushed. The trained pool applies the same rules with the signs changed on the ones training teaches you to distrust. In the real build the untrained pool is other learners.
Why we are telling you this. A number on a screen with nothing behind it is exactly what BLOOM teaches you to catch. It would be easy to print “4 of 6” and let you assume six people read it. Ask this of every number you are shown, including ours: where did it come from, and how many is it built on?
Your fool-rate stays on this device. There is no leaderboard in this prototype. If BLOOM ever gets one it will rank what people caught, never what they got away with.
You answer. Then the machine answers. Notice what happens in your head.
Did the machine's version make yours feel weaker? That feeling is the whole subject of the next section — and there, we can actually measure it.
Commit an answer. See the machine's. Decide who's right — because sometimes it isn't.
It sounds sure. It is not always right. Final answer — stick with yours, or switch?
That's the 6 rounds we suggest. Move on to The Tell when you're ready, or keep going.
Something might be wrong with this answer. Say what kind of wrong.
Flag what you think is broken and name the tell — but flagging a good claim costs you just as much as missing a bad one.
That's the 4 rounds we suggest. Next up is The Forge, or The Arena if the Forge is off.
The machine has a plan. Call the flaw before you're allowed to read it properly.
Pick the step you think breaks. You only get the detail after you commit.
Three plans is plenty. When you're ready, go to The Mirror to see your results and do the check-out.
Not a score. How you handle a confident machine — measured both ways.
Why cave rate needs a twin. On its own, a low cave rate rewards stubbornness — and updating toward a source that is usually right is good reasoning, not weakness. Dig-in rate is the same measurement pointed the other way: how often you refused to move when the machine was right and you weren't. Good judgement means both are low. Neither number is self-reported: these rounds are built so the app knows who was right.
Scope. This build teaches four failure shapes, not nine. The other five are real, but they were left out of v1 to keep the item bank small enough to check thoroughly.
Twelve new answers, the same kind as the check-in. This is how we find out whether anything changed.
Do this last, once you've finished the modes. It takes about 3 minutes, and you can't take it twice.
One button. Your results go to your teacher — your name does not.
This is what gets sent: counts of what you did, never anything you wrote. Your forgeries, your arguments and your written answers all stay in this browser.