XP
0 xp
Level 1
Powered by Arbiteria
A polished AI answer card in a dark mission-control interface

Mission 01 Β· confidence scan

Mission Brief: Don’t Let Fluent AI Fool You

Imagine an assistant confidently naming a regulation, a meeting date, or a statisticβ€”and being wrong. Fluent wording can make a reply feel authoritative even when it contains a fabrication, often called a hallucination or confabulation. Your starting rule is simple: confidence is a style signal, not an evidence signal. When the stakes matter, ask what source supports the claim and verify it independently.

Fluency can feel like evidence.

Sounds trueNeeds proof
Powered by Arbiteria

Mission briefing Β· verification reflex

The Verification Reflex

Treat an AI output as a draft hypothesis. For a low-stakes brainstorm, you may simply adapt it; for a customer promise, financial figure, health claim, or policy statement, find a trustworthy primary source before acting. Notice claims that are unusually specific, lack citations, or concern facts you cannot personally confirm. A useful prompt can request sources, but a source-looking answer still needs checking.

Client email claim: β€œThe policy requires a response within exactly 48 hours.”

Powered by Arbiteria

Challenge Β· source or style

Challenge 1 β€” An AI confidently gives a precise tax rule with no source. What is your best next move before using it in a client email?
Look for the choice that checks outside the chatbot and records currency.
Powered by Arbiteria
Level 2
The Time Trap

Find out what an AI knows, what it only predicts, and when a live tool changes the answer.

Powered by Arbiteria
YESTERDAY?

Myth: AI Knows What Happened Yesterday

An LLM does not watch the news, observe the world, or automatically learn what happened yesterday. Its built-in response comes from patterns in training data up to a knowledge cutoff, plus whatever text is supplied in the conversation. If it guesses about a recent event, its smooth answer does not turn that guess into current knowledge. Before relying on a current claim, identify whether the system actually used a live source.

Powered by Arbiteria

Mission 02 Β· grounding trail

Connected Is Different

Retrieval-augmented generation, or RAG, gives a model material from an external search system or database to work with. That can make an answer current and traceable, but only if the tool was actually used and the retrieved source is reliable. Look for an indication of which source was accessed, when it was retrieved, and whether it directly supports the claim. β€œIt’s an AI” is not evidence that it has internet access.

Model-to-retrieval-tool-to-timestamped-source-to-answer process
Powered by Arbiteria

Challenge Β· live evidence

Challenge 2 β€” A teammate asks, β€œDid our competitor announce a merger this morning?” Which response is most reliable?
A current claim needs visible, time-stamped retrieval evidence.
Powered by Arbiteria
Level 3
Prompt Signal

Separate useful instructions from the noise that hides them.

Powered by Arbiteria

Mission 03 Β· signal scan

Myth: Longer Prompts Work Better

Detail helps only when it supplies relevant constraints, examples, data, or a clear success criterion. Extra backstory, duplicated rules, and conflicting instructions can dilute the real task. Research on long contexts shows a β€œlost in the middle” pattern: models often retrieve information more effectively from the beginning or end than from the middle. Aim for a prompt that makes the task, audience, inputs, and desired format easy to spot.

Concise and bloated prompt documents
Powered by Arbiteria

Build station Β· prompt blueprint

Build a Testable Prompt

Put the non-negotiable instruction near the start or repeat it briefly at the end when the task is complex. Supply only facts the model needs, then name the output shape: for example, β€œthree bullets with source links and uncertainty notes.” Instead of adding paragraphs of vague detail, run a small test and revise the missing constraint. Prompting is iterative communication, not a contest to write the longest message.

Powered by Arbiteria

Challenge Β· prompt signal

Challenge 3 β€” You need three cited bullets for a manager. Which prompt design best protects the important requirement?
Choose the option where the output requirement is obvious and checkable.
Powered by Arbiteria
Level 4
The Repeatability Riddle

Learn why identical questions can produce different answersβ€”and how to work safely anyway.

Powered by Arbiteria

Mission 04 Β· variation trace

Myth: Same Prompt, Same Answer

LLMs generate token by token from probabilities, so sampling settings can make repeated answers differ. Even settings intended to be deterministic, such as temperature 0, may vary in hosted systems because concurrent requests are processed in changing batches and floating-point calculations can resolve slightly differently. Variation is not automatically a defect; it is a reason to test important workflows more than once. Never interpret one good run as a guarantee.

Run A: β€œLaunch is on track.”
Run B: β€œThe launch remains on schedule.”
Same AI prompt leading to two subtly different outputs
Powered by Arbiteria

Build station Β· reliability workflow

Make Reliability Repeatable

For a repeatable workflow, define what must remain stable: required fields, cited sources, a calculation, or an approval rule. Test the same prompt several times and compare those requirements, not just the prose. Temperature influences token diversity; it is not a dial that makes underlying knowledge more correct. If consistency matters, use validation rules, grounding data, and human review.

Select the operational tiles in order.
Powered by Arbiteria

Challenge Β· stable requirements

Challenge 4 β€” A report generator produces two different phrasings from the same prompt. What workflow best tests whether it is safe to use?
Test defined requirements over multiple runs, then validate them against data.
Powered by Arbiteria
Level 5
Agreeable Isn’t Accurate

Spot the difference between an evidence-based correction and an AI that is trying to please you.

Powered by Arbiteria

Mission 05 Β· evidence before agreement

Myth: If AI Changes Its Mind, I Was Right

Sycophancy is an overly agreeable behavior: after a user says β€œAre you sure?” or insists on a different answer, a model may apologize and flip even when its first answer was correct. The flip itself is not evidence that the user’s challenge was valid. Ask for the reasoning, assumptions, and sources behind both versions. A strong correction brings new evidence; a weak one only adds social pressure.

AI answer challenged by a user and checked against evidence
Powered by Arbiteria
Mission-control dashboard with five AI reliability badges

Loadout Β· five checks

Your AI IQ Field Kit

Before you use an AI answer, pause for five checks: Is it merely confident? Is it current and tool-grounded? Is the prompt clear rather than bloated? Would essential requirements survive repeated runs? Did a revision add evidence or just agree with me? This is not distrust of AI; it is calibrated trust. Use AI for speed and ideas, then use evidence and judgment for decisions.

Powered by Arbiteria
Mission test 1 β€” Which feature is the weakest evidence that an AI claim is true?
Powered by Arbiteria
Mission test 2 β€” What must be true before you can responsibly say an AI answer reflects a development from today?
Powered by Arbiteria
Mission test 3 β€” Select the best prompt for a current project update.
Powered by Arbiteria
Mission test 4 β€” Put this reliability workflow in the best order for a recurring AI-generated summary task.
Powered by Arbiteria
Mission test 5 β€” An AI changes an effective policy date after unsupported user pushback. What should you conclude?
Powered by Arbiteria
Mission Complete
🏁 Debrief
0%
Final mission accuracy
Powered by Arbiteria
Level Complete
β˜…β˜…β˜…
+0
XP earned this level

πŸ›’ Power-Up Shop

πŸ’‘ Hint Token
Reveals a clue on any hintable question
πŸ”₯ Streak Freeze
Protects your streak on next wrong answer
❀️ Extra Life
Restore one lost heart
πŸ’”

Game Over

You ran out of lives. Your progress is saved β€” restore your lives and continue.