Reasoning First, and the Weakness of Negatives
Ask for Working You Can Check
On problems with several steps — comparisons, calculations, anything where a conclusion depends on a chain of logic — asking the model to show its working generally improves the result. The mechanism is not mysterious: the intermediate steps become part of the text the model is conditioning on, so the conclusion is drawn from worked reasoning rather than produced cold. Frontier models increasingly reason before answering by default, sometimes invisibly — so the deeper reason to ask is not to cause thinking but to get an account you can check: the numbers, the assumptions, the evidence a conclusion rests on. You can read the working. A wrong answer with visible steps is debuggable; a wrong answer on its own is just a wrong answer.
- Weak: "Which of these three options is cheapest overall?"
- Strong: "Work out the three-year total cost for each option, showing the numbers, then state which is cheapest and why"
- Visible reasoning lets you find where it went wrong instead of just rejecting the answer
- Read the steps — a confident conclusion sitting on a flawed step is the main risk
Structured Thinking Beats "Think Step by Step"
The generic instruction helps, but naming the actual steps helps more, because you know your problem and the model does not. For a hiring decision: "First list the requirements. Then score each candidate against each requirement with a one-line justification. Then note where the evidence is thin. Only then give a recommendation." For a bug: "First restate what the code is supposed to do, then trace what it actually does with this input, then identify the first point where they diverge." You are supplying the method, not just asking for effort. This is also where a lot of professional expertise lives — the steps you take are the thing you know that a general-purpose model does not.
- Name the actual steps of your method rather than asking generically for reasoning
- Insert a step for "where is the evidence weakest?" before the conclusion
- Ask for the conclusion last, explicitly, so it does not lead the reasoning
- Your professional method is valuable prompt content — write it down once and reuse it
Why "Do Not" Is Weaker Than "Do"
Negative instructions work, but less reliably than positive ones, and it is worth understanding why. A negative tells the model what to avoid without saying what to do instead, which leaves the space of acceptable answers wide open. It also puts the unwanted thing in the context, where it exerts some pull. "Do not use bullet points" is weaker than "write this as three flowing paragraphs". "Do not be too formal" is weaker than "write it the way you would message a colleague you like". The practical rule: whenever you catch yourself writing a "do not", ask what you want instead and write that. Keep the negative too if it is a hard boundary, but never let it be the only instruction.
- Weak: "Do not make it sound corporate." Strong: "Write it the way you would explain it to a friend over coffee."
- Weak: "Do not include a conclusion." Strong: "End on the last recommendation, with no summary section."
- A negative alone leaves everything else permitted; a positive narrows the target
- Keep hard prohibitions, but always pair them with the desired behaviour
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.