Powered by Arbiteria
Mission 01 Β· confidence scan
Imagine an assistant confidently naming a regulation, a meeting date, or a statisticβand being wrong. Fluent wording can make a reply feel authoritative even when it contains a fabrication, often called a hallucination or confabulation. Your starting rule is simple: confidence is a style signal, not an evidence signal. When the stakes matter, ask what source supports the claim and verify it independently.
Powered by ArbiteriaMission briefing Β· verification reflex
Treat an AI output as a draft hypothesis. For a low-stakes brainstorm, you may simply adapt it; for a customer promise, financial figure, health claim, or policy statement, find a trustworthy primary source before acting. Notice claims that are unusually specific, lack citations, or concern facts you cannot personally confirm. A useful prompt can request sources, but a source-looking answer still needs checking.
Powered by ArbiteriaChallenge Β· source or style
Powered by Arbiteria
Powered by ArbiteriaAn LLM does not watch the news, observe the world, or automatically learn what happened yesterday. Its built-in response comes from patterns in training data up to a knowledge cutoff, plus whatever text is supplied in the conversation. If it guesses about a recent event, its smooth answer does not turn that guess into current knowledge. Before relying on a current claim, identify whether the system actually used a live source.
Powered by ArbiteriaMission 02 Β· grounding trail
Retrieval-augmented generation, or RAG, gives a model material from an external search system or database to work with. That can make an answer current and traceable, but only if the tool was actually used and the retrieved source is reliable. Look for an indication of which source was accessed, when it was retrieved, and whether it directly supports the claim. βItβs an AIβ is not evidence that it has internet access.

Powered by ArbiteriaChallenge Β· live evidence
Powered by Arbiteria
Powered by ArbiteriaMission 03 Β· signal scan
Detail helps only when it supplies relevant constraints, examples, data, or a clear success criterion. Extra backstory, duplicated rules, and conflicting instructions can dilute the real task. Research on long contexts shows a βlost in the middleβ pattern: models often retrieve information more effectively from the beginning or end than from the middle. Aim for a prompt that makes the task, audience, inputs, and desired format easy to spot.

Powered by ArbiteriaBuild station Β· prompt blueprint
Put the non-negotiable instruction near the start or repeat it briefly at the end when the task is complex. Supply only facts the model needs, then name the output shape: for example, βthree bullets with source links and uncertainty notes.β Instead of adding paragraphs of vague detail, run a small test and revise the missing constraint. Prompting is iterative communication, not a contest to write the longest message.
Powered by ArbiteriaChallenge Β· prompt signal
Powered by Arbiteria
Powered by ArbiteriaMission 04 Β· variation trace
LLMs generate token by token from probabilities, so sampling settings can make repeated answers differ. Even settings intended to be deterministic, such as temperature 0, may vary in hosted systems because concurrent requests are processed in changing batches and floating-point calculations can resolve slightly differently. Variation is not automatically a defect; it is a reason to test important workflows more than once. Never interpret one good run as a guarantee.

Powered by ArbiteriaBuild station Β· reliability workflow
For a repeatable workflow, define what must remain stable: required fields, cited sources, a calculation, or an approval rule. Test the same prompt several times and compare those requirements, not just the prose. Temperature influences token diversity; it is not a dial that makes underlying knowledge more correct. If consistency matters, use validation rules, grounding data, and human review.
Powered by ArbiteriaChallenge Β· stable requirements
Powered by Arbiteria
Powered by ArbiteriaMission 05 Β· evidence before agreement
Sycophancy is an overly agreeable behavior: after a user says βAre you sure?β or insists on a different answer, a model may apologize and flip even when its first answer was correct. The flip itself is not evidence that the userβs challenge was valid. Ask for the reasoning, assumptions, and sources behind both versions. A strong correction brings new evidence; a weak one only adds social pressure.

Powered by Arbiteria
Loadout Β· five checks
Before you use an AI answer, pause for five checks: Is it merely confident? Is it current and tool-grounded? Is the prompt clear rather than bloated? Would essential requirements survive repeated runs? Did a revision add evidence or just agree with me? This is not distrust of AI; it is calibrated trust. Use AI for speed and ideas, then use evidence and judgment for decisions.
Powered by Arbiteria
Powered by Arbiteria
Powered by Arbiteria
Powered by Arbiteria
Powered by Arbiteria
Powered by Arbiteria
Powered by ArbiteriaYou ran out of lives. Your progress is saved β restore your lives and continue.