A strange question that works
Someone once asked: why does cannabis seem to make people more creative?
The honest answer is that it does not make anyone smarter. What it changes is filtering. Under its influence, the mind lets through associations it would normally reject — the same “leaky” attention that shows up in people with high creative achievement, whose latent inhibition is measurably lower (Carson et al., Journal of Personality and Social Psychology, 2003). Distant ideas reach awareness more easily. That feels like creativity from the inside.
Then someone asked a better question: so it’s like turning up a model’s temperature?
Yes. And that comparison opens a door. Every sampling parameter in a language model sits on the same trade-off — how boldly the model picks its next word — and each one has a rough equivalent in how a mind works. That is the pharmacopoeia this article draws. One caveat before we start, because it is also the most interesting thing here: the dial is disappearing. The newest reasoning models expose no temperature at all. The industry stopped turning the knob up and started spending the compute on thinking instead — which is this article’s argument, arrived at from the other direction.
Read this first. These are analogies, not neuroscience. THC does not set a dial in the brain, and no substance has a
top_pvalue. A model is not a mind whose filters got loose — it has no subjective state to lose. The comparison earns its place because a person and a model both face a version of the same decision: choosing among thousands of possible next steps, and deciding how much to trust the obvious one.
Be precise about where it holds. A model makes one decision, from a probability distribution. A substance intervenes in a whole brain, with memory, mood, and metabolism attached. Everything below lives in the narrow strip where those two overlap.
The one trade-off behind everything
A language model does not “know” what to say. At every step, it produces a probability for every token it could write next, then picks one. A token is not quite a word — it can be a whole word, but it can equally be a fragment with a leading space — and that difference matters later.
Two things follow from that:
- Pick the most likely word every time and the text is safe, correct, and boring. It loops.
- Pick something unlikely often enough and the text is surprising — sometimes brilliant, sometimes nonsense.
Every parameter below just moves the needle between those two poles.
The LLM pharmacopoeia
Each row is one dial. The two right-hand columns show the same parameter turned low and high — the size, not a different setting.
⚡ stimulant · 🌿 THC · 🌀 psychedelic · 🍷 alcohol · 💤 benzodiazepine · 💊 opioid
Treat these six as poles, not owners. The substance named at each end of a row is the state that end resembles, and the same one can appear under two different dials — because the dials share one axis: how wide the filter is. Temperature is the clearest case; the others borrow its vocabulary.
| Parameter | Turn it down ↓ | Turn it up ↑ |
|---|---|---|
| Temperature | Tight focus — ⚡ stimulant: locked in, predictable, the same answer every time | Loose filtering — 🌿 THC: unusual ideas surface, and so does noise |
| Top-p | Narrow net — 💤 benzodiazepine: only the obvious words qualify | Wide net — 🌀 psychedelic: odd candidates stay on the table |
| Context window | Drunk short-sightedness — 🍷 alcohol: locally fluent, forgets why | Clear working memory — ⚡ stimulant: holds the whole story |
| Repetition penalty | Comfortable habit — 💤 benzodiazepine: stable wording stuck in its groove | Restlessness — ⚡ stimulant: fresh phrasing, sometimes forced |
| Max tokens | Cut off mid-thought — 💊 opioid: the thought starts and never lands | Room to finish — ⚡ stimulant: the stamina to carry one thought to its end |
Max tokens is the odd one out here, and it earns its own section below. Strictly speaking it is not a sampling parameter at all: it never touches which token gets picked, only how many are allowed. It is a budget, not a mood.
Temperature: the THC of the model
This is the comparison that started it all, and it is the cleanest one.
Temperature scales the probabilities before the model picks. A low value sharpens them: the favourite word becomes overwhelming, and the model almost always takes it. A high value flattens them: the long shots become real contenders.
| Temperature | The model’s behaviour | The human parallel |
|---|---|---|
| 0.0 – 0.3 | Nearly deterministic. Best for code, facts, formatting. | Locked in, narrow, reliable — the ⚡ stimulant focus |
| 0.7 – 0.8 | Balanced. Everyday conversation and writing. | Engaged and flexible — the ⚡ mild stimulant baseline: a coffee, not a trip |
| 1.0 – 1.2 | Colourful, associative, takes risks. | Loose and creative — 🌿 THC |
| > 1.5 | Coherence collapses. Words stop fitting together. | Too far gone to judge your own ideas — 🌀 psychedelic territory |
That last row is the important one, and there is real evidence behind it. There is a ceiling. Past a certain point you do not get more creativity, you get noise. In regular cannabis users, high-potency cannabis measurably impairs divergent thinking, while the low-potency version changes nothing (Kowal et al., Psychopharmacology, 2015) — the ceiling is not a metaphor, it shows up in the lab.
A second finding matters just as much. When acute cannabis does raise divergent thinking, it does so mainly in people who start low, not in people who are already creative (Schafer et al., Conscious Cognition, 2012). So the honest claim is not “a drug makes you creative”. It is that loosening the filter moves you off your baseline — in whichever direction your baseline was not.
A third result cuts against the advice this article is heading toward, so it belongs here rather than in a footnote. In a systematic study of nine models across five prompt-engineering techniques, moving temperature anywhere from 0.0 to 1.0 made no statistically significant difference to problem-solving performance (Renze & Guven, Findings of EMNLP, 2024). Temperature moves what the model says; it does not reliably move whether it is right. Two caveats follow from that. Divergent thinking and verbal fluency — the lab’s yardsticks above — are contested proxies for creativity rather than creativity itself; and the effect sizes are small enough that “turn it down” is a habit, not a law.
Where the analogy breaks: the person loses the faculty that would catch a bad idea. The model never had one. There is no inner critic built into a transformer for a high temperature to overwhelm — there is a distribution, and the model samples from it. Any critic has to be added from the outside, which is exactly what best-of-n and self-consistency do further down. The person’s failure is a lost faculty; the model’s is an absent one. Same visible result, different mechanism, and the difference is worth keeping straight.
Top-p: how wide the net is
Temperature changes how sharply the model prefers one word. Top-p changes how many candidates it is allowed to consider at all.
With top_p = 0.9, the model takes the smallest set of words whose combined probability reaches 90%, and ignores everything outside it. Lower it to 0.5, and it only ever considers the obvious half. Raise it to 1.0, and even very unlikely words stay on the table.
Think of it as how open you are to strange suggestions. A narrow net is the 💤 benzodiazepine conversation — calm, expected, nothing surprising. A wide net is the 🌀 psychedelic one — where the strange thought is allowed to arrive.
Temperature and top-p overlap, so in practice you usually tune one and leave the other near default. Changing both at once is like mixing two substances and trying to guess which one caused what.
Context window: memory, not creativity
This is where the analogy gets more interesting than temperature.
The context window is everything the model can see at once — the prompt, the conversation, the documents you pasted in. It is not about picking words; it is about holding the thread.
Imagine writing a novel. With full context, the model remembers the characters, their motives, the rules of the world, and the decisions made in chapter two. Cut the context to a fifth and it still writes beautiful sentences — but it starts to forget why the hero is angry. The prose stays good; the story stops adding up.
That is the state alcohol produces: locally fluent, globally lost. The person still talks well. They just cannot hold the whole picture, and — crucially — they do not notice the gap.
The lesson is that big context windows are not just “more room”. They are what lets a model check a new idea against everything that came before.
Repetition penalties: breaking the groove
These parameters push the model away from words it has already used.
- Repetition penalty — a flat penalty on any repeated token.
- Presence penalty — a one-time nudge once a word appears at all.
- Frequency penalty — a growing penalty the more a word appears.
Turned up, they produce restlessness: fresh phrasing, new angles, an unwillingness to settle. Turned up too far, they produce the opposite problem — the model avoids the right word only because it already used it, and the text starts to strain.
That is the familiar human pattern of desperately avoiding a rut. It is the ⚡ stimulant itch — needing the next word to be different — and pushed too far it turns compulsive, the way a 🌀 psychedelic loop keeps circling a phrase that has lost its meaning.
Max tokens: the length of the rope
Every parameter so far changes how the model chooses. Max tokens changes only how long it is allowed to keep choosing. It is a hard cap on the output: when the count runs out, generation stops at the next token boundary. Since tokens are chunks rather than words, that can leave you without the final word of a sentence — but never with half a word on the page.
That makes it the odd one out, and the substances show why. Here is the same question — a little, or a lot — asked of each one:
| Substance | Max tokens ↓ — a little | Max tokens ↑ — a lot |
|---|---|---|
| 🌿 Cannabis / THC | The thought stops on its own — it loses the thread rather than hitting a wall | More room, but the thread wanders sideways instead of growing longer |
| 🌀 Psychedelics | Rarely cut off — time stretches instead of ending | A much longer subjective duration, but not a longer argument |
| ⚡ Stimulants | Cut off while still winding up | The stamina to carry one thought all the way to its end |
| 🍷 Alcohol | Blackout mid-sentence — the ending is simply gone | The long, rambling story that finally reaches the point |
| 💤 Benzodiazepines | Gently trails off into sleep, mid-thought | More runway — but sedation still ends it before the rope does |
| 💊 Opioids | The nod: the sentence dissolves before it lands | More words available, but the drift is the same — length is not the problem |
Look at the pattern. The “a little” column is vivid for every substance — being cut off is a real, recognisable state, and each one names it differently. The “a lot” column describes something much thinner: more room, a longer stretch, and in exactly one row — the stimulant — a genuinely better state, the stamina to finish the thought.
That one row is the honest exception, and it is worth naming rather than hiding. Finishing an argument is a capability. But note what it depends on: whether more tokens help is a property of the task, not of the model. If the thought fits in the budget, a bigger budget buys nothing; if it does not, the bigger budget is the difference between an answer and a truncation. Nothing about the model changes either way.
That is the tell. Most parameters are moods: they change what the model is like. Max tokens is a fence, and the only thing it can do wrong is run out too early.
The parameters that are not about mood
Max tokens exposed the limit of the method: it has a low and a high, but no state of mind to match them. Three settings never even pretend to have one — the closest substance is listed anyway, because that is where the analogy visibly gives up. Note that it also reuses substances claimed by dials above: the same pole, seen from another direction.
| Parameter | What it actually does | Nearest substance comparison | Why it still breaks |
|---|---|---|---|
| Seed | Fixes the random draw | 🌿 THC — same strain, same dose, same room | The same night never comes twice: tolerance and mood shift the trip — and on a hosted API, “same seed” is a best-effort promise, not a guarantee |
| Stop sequences | Ends generation at a marker | 🍷 alcohol — last call: the bartender ends the round | A rule imposed from outside, not a state you reach |
| System prompt | Standing instructions before the task | 🌀 psychedelic — set and setting: the room and the people | It sets what the model is being, not how boldly it picks |
Three more dials the pharmacopoeia skips
The comparison is not exhaustive. Three standard controls sit outside it, and each earns a line precisely because it shows the same axis from a different angle:
| Parameter | What it actually does | Why there is no substance |
|---|---|---|
top_k | Keeps only the k most likely tokens, however flat or peaked the distribution | A fixed headcount, not a mood — the oldest and crudest version of “how wide is the net” |
| Min-p / typical sampling | Truncates by a relative threshold: keep tokens whose probability is at least a fraction of the top token’s | Newer, and better behaved at high temperature (Nguyen et al., ICLR, 2025) — the same idea as top-p, tuned differently |
| Beam search | Not sampling at all: keeps several candidate sequences alive and returns the most probable whole one | Deterministic, fluent, and famously bland — the sober colleague who writes the meeting minutes |
The part everyone forgets
Here is the insight that makes the whole analogy worth writing down.
More unusual associations do not mean better creativity.
A high-temperature model generates surprising text. A person on THC feels like they are having brilliant ideas. Both are genuinely producing more novelty — and both are worse at judging whether that novelty is good.
The model cannot reliably rank its own wilder output. The person, high, loses exactly the faculty needed to tell a great idea from a strange one. The ideas still look great from the inside. That is the trap: the feeling of insight survives; the quality control does not.
And the feeling is not a small thing to dismiss — but it is a poor guide. Sober cannabis users report higher creativity and score better on one objective measure, yet both advantages disappear once you control for openness to experience (LaFrance & Cuttler, Conscious Cognition, 2017). What looks like the drug’s gift is a personality trait the user brought with them. A confidence signal is not an accuracy signal, in a person or in a model.
This is where the analogy earns its keep, because the fix is the same in both cases and it is not “turn the dial back down”. It is to separate generation from judgement. Best-of-n sampling draws many candidates at a high temperature and lets a separate, cooler pass pick the winner. Self-consistency does the same with answers instead of prose. The point of both is that novelty is only worth having when something clean is around to evaluate it — which is also the argument for sleeping on an idea, or showing your work to someone sober.
This is why the useful direction is almost always down, not up. Lower the temperature when the answer must be correct. Tighten the top-p when you need the conventional answer. Give the model more context when the task is long. Creativity in a system — silicon or biological — is not mainly about generating more; it is about keeping the good ideas and dropping the rest.
The cheat sheet
| If you want… | Reach for | Substance parallel |
|---|---|---|
| Correct, repeatable answers | temperature → 0 | ⚡ stimulant — locked in, sober, precise |
| Natural prose and conversation | temperature → 0.7 | ⚡ mild stimulant — a coffee, not a trip |
| Unusual ideas for a brainstorm | temperature 1.0–1.2, top_p 0.95–1.0 | 🌿 THC + 🌀 psychedelic — loose and associative |
| A model that keeps the whole story | Bigger context window | ⚡ stimulant — clear working memory |
| Fresh wording, no loops | repetition_penalty 1.1–1.3, or presence_penalty 0.1–0.5 | ⚡ stimulant — restless, avoids old grooves |
| Consistent output across runs | Fixed seed (best-effort on hosted APIs) | 🌿 THC — same strain, same dose |
In practice that looks like {"temperature": 0} for extraction and classification, {"temperature": 0.7} for conversation, and {"temperature": 1.0, "top_p": 0.95} for a brainstorm you will filter afterwards. Move one dial at a time — changing temperature and top_p together, as the top-p section warns, makes it impossible to tell which one did the work.
And when you move to a reasoning model, the dial has not vanished — it changed shape. Instead of temperature you get reasoning_effort (low / medium / high, or the provider’s equivalent): how hard the model thinks before it answers. That is a thinking dial standing where the sampling dial used to be. The knob did not disappear; it moved from the output to the deliberation.
What the analogy is really for
The comparison earns its place because it forces one idea into the open: intelligence is not only about producing options. It is about choosing among them.
A model with the temperature cranked up is not more intelligent than a careful one. A person with their filters loosened is not more creative in any way that survives the morning. Both systems, when you push the dial far enough, trade judgement for novelty and call the result insight.
So use the pharmacopoeia. Reach for a lower temperature when the stakes are correctness, and a higher one when you are genuinely exploring. Just remember which direction actually makes the work better — and that no parameter, and no substance, has ever done the thinking for you.
If you want to go deeper: the parameters above are the standard sampling controls in most model APIs. The names are consistent, the defaults are not — always check the docs for the model you are using. And on the newest reasoning models there is often nothing to check: temperature is not exposed at all, and what used to be max_tokens is now a length budget under another name — max_completion_tokens on OpenAI’s Chat API, max_new_tokens in Hugging Face Transformers. The name moved; the dial did not.
References
- Carson SH, Peterson JB, Higgins DM., Decreased latent inhibition is associated with increased creative achievement in high-functioning individuals, Journal of Personality and Social Psychology 85(3), 2003 — pubmed.ncbi.nlm.nih.gov/14498785
- Kowal MA et al., Cannabis and creativity: highly potent cannabis impairs divergent thinking in regular cannabis users, Psychopharmacology 232(6), 2015 — pubmed.ncbi.nlm.nih.gov/25288512
- Schafer G et al., Investigating the interaction between schizotypy, divergent thinking and cannabis use, Conscious Cognition 21(1), 2012 — pubmed.ncbi.nlm.nih.gov/22230356
- LaFrance EM, Cuttler C., Inspired by Mary Jane? Mechanisms underlying enhanced creativity in cannabis users, Conscious Cognition 56, 2017 — pubmed.ncbi.nlm.nih.gov/29065317
- Renze M, Guven E., The Effect of Sampling Temperature on Problem Solving in Large Language Models, Findings of EMNLP 2024 — aclanthology.org/2024.findings-emnlp.432
- Nguyen MN, Baker A, Neo C, Roush A, Kirsch A, Shwartz-Ziv R., Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs, ICLR 2025 — arxiv.org/abs/2407.01082
- OpenAI API reference, Create chat completion —
max_tokensdeprecated in favour ofmax_completion_tokens; reasoning models document unsupported sampling parameters — developers.openai.com/api/reference/resources/chat - Hugging Face Transformers, Generation strategies — decoding strategies and
max_new_tokenssemantics — huggingface.co/docs/transformers/en/generation_strategies