Home / Blog / How to Make Claude Work: One Chat, One Task
Guide

How to Make Claude Work: One Chat, One Task

20 August 2026 · John Carlsson

1. What I'm actually claiming

Let me put it plainly so there's no wriggle room later.

The claim is this: the more you load into a single chat, the worse Claude performs. Not vaguely worse. Not occasionally worse. Reliably, measurably worse — and it gets worse the more you pile in. "Load into a chat" means all of it: a long back-and-forth conversation, big documents you've pasted, files, and yes, stacks of so-called magic prompts. It's all the same weight to the machine.

From that one fact, everything else follows. If loading it up is what breaks it, then the fix cannot be to load it up with cleverer instructions. The fix has to be the opposite: keep each chat small and clean. One task, finish it, start again. That's the whole thing. The rest of this document is me showing you it's true, because you shouldn't take it on my say-so any more than on some prompt-seller's.

---

2. How Claude actually works — and why nearly everyone gets this wrong

Most people picture Claude like a person you're talking to. You said something earlier, so it "remembers" it, the way a colleague would. That picture is wrong, and the wrongness is the root of the whole problem.

Claude has no memory of the conversation as it goes along. Between one message and the next it holds nothing on its own. It doesn't sit there carrying the thread in its head. This is what people in the field mean when they call these systems "stateless" — state, meaning the running memory of what's happened, simply isn't kept.

So how does it seem to follow a conversation at all? Because every single time you press send, the entire chat so far — every message from you, every reply it gave, every document you pasted, every file — is gathered up and fed back into it from cold. It reads the whole pile again, from the top, every turn, and only then writes the next reply. It is not continuing a conversation. It is re-reading the entire transcript from scratch, over and over, and each time producing the next line.

Hold onto that image, because it explains everything: every turn, it re-reads everything.

In a short chat this costs nothing. There's barely anything in the pile, so re-reading it is easy and it stays bang on the point.

But watch what happens as the chat grows.

---

3. Why a growing chat rots the answers

As the conversation goes on, that pile it re-reads every turn gets bigger and bigger. And now two things go wrong at once.

First: it can't reliably tell what still matters from what's finished with. An instruction you gave and dealt with an hour ago is still sitting in the pile. A file you were working on and moved past is still in the pile. A correction you made, agreed, and left behind — still in the pile. To Claude, re-reading the whole lot fresh each turn, that old finished stuff pulls on the answer just as hard as the thing you asked ten seconds ago. It has no reliable sense of "that's done, ignore it." So it reaches in and grabs the wrong bit — an old order, a job already completed, a mistake you already corrected — and drops it into the reply. That is exactly the daft error you've seen a hundred times: it does the thing you told it *not* to, or redoes something already done, or drags back a detail you'd killed off.

Second, and worse: the mistakes compound. The moment it takes one wrong turn, that wrong turn is now part of the pile too. Next turn, it re-reads its own error along with everything else — and builds on it. Nothing gets dropped. So a single early slip doesn't stay a single slip; it becomes the foundation the next few answers are stacked on. This is why a long chat doesn't just make *occasional* mistakes — it spirals. One wrong assumption early and half an hour later you're arguing with something that's built three more things on top of it.

There's a solid, boring reason underneath this, and it's worth knowing because it kills the "just prompt it better" nonsense stone dead.

These models work on what researchers describe as a fixed "attention budget" — a limited pot of focus that has to be spread across everything in front of them. Think of it exactly like a person's concentration. Ask someone to hold one instruction in their head — easy, full focus on it. Ask them to hold two hundred at once — they start dropping things, mixing them up, losing the important one under the trivial ones. The pot of focus didn't grow; it just got divided into thinner and thinner slices. Claude is the same. Every extra token you add — every extra sentence, document, prompt — is one more thing the same fixed focus has to be split across. Add enough and the bit that actually matters gets such a thin slice of attention that it effectively drowns. More in means less focus on each thing, which means more of the important things dropped. That's not a flaw you can prompt your way around. It's the shape of the machine.

---

4. Why the memory setting quietly makes it worse

There's a second load most people never think about.

Claude keeps saved notes between chats — a little memory of facts about you and your work, so a fresh chat knows who you are without you re-explaining every time. Used properly, that's a good thing. It's how a new chat can start already knowing the basics.

But here's the catch. Those saved notes get loaded into the pile at the very start of every single chat — before you've typed one word. So the state of those notes decides how clean your fresh start actually is.

Keep them short — a handful of real facts — and a new chat starts light, as it should. But let them grow fat, or (the classic mistake) paste a whole long handover document into them, and now every new chat starts already part-loaded with all that weight. You've thrown away the clean start before you've even begun. The attention budget is already being split before your first question lands. And it makes the mistakes from section 3 show up straight away, in a chat that ought to have been fresh.

This is the sting in the tail: a bloated memory doesn't drag on one chat. It drags on every chat you ever open, permanently, until you trim it. Which is exactly why the memory has to be kept lean on purpose — a few facts, never whole documents.

---

5. Why "here's 7 prompts for Claude" is garbage — the bit that started all this

Now the thing that set me off, and by this point the argument writes itself.

Everywhere you look, someone is selling a stack of clever prompts. "Paste these seven magic instructions at the start and Claude transforms." There are whole threads, whole products, built on it.

Put it against everything above and it collapses on contact. A stack of prompts is just more stuff crammed into the pile. It is more for that fixed attention budget to be split across — before you've even asked your real question. You are not tuning the machine up. You are pre-loading it with clutter and starting it off already half-swamped. Those seven prompts don't sharpen it; they blunt it, for the exact same reason a long chat blunts it. The sellers have it precisely backwards: they're handing you more load and calling it a fix, when load is the disease.

Now, the honest part — because I'm not going to overclaim just to win the point. A good, plain instruction genuinely does work. "Keep it short." "Don't guess — ask me." "Show me proof it worked." Those help, because they're small and clear and they cost almost nothing in the budget. The point was never that instructions are worthless. The point is two things. One: one or two clear instructions do the real work — you don't need seven, and the seventh is doing more harm than the first is doing good. Two: no instruction, however clever, rescues a chat that's already overloaded — once the pile is too big, prompting is rearranging deckchairs.

So the mega-prompt sellers are flogging you a cure that is, ingredient for ingredient, a dose of the disease. It only *looks* like expertise because it's complicated. The real answer is almost insultingly simple, which is probably why nobody can sell it: one chat, one task.

---

6. The proof — because you should not take my word for it either

I've just spent five sections telling you how it works. Fair enough for you to say: prove it. Two ways — one you can do yourself in three minutes, one from the people who measured it properly.

6.1 Run it yourself

1. Open a brand-new chat. Give Claude a small, clear task. Watch it handle it cleanly and correctly. 2. Go to a long chat you've had running for ages, stuffed with back-and-forth. Give it the exact same task. 3. Watch the mistakes creep in — it forgets part of it, grabs the wrong thing, drifts back to something you sorted earlier.

Same model. Same task. The only variable you changed is how much was already loaded in. Whatever difference you see is caused by that and nothing else. That's not a demo I'm controlling — it's an experiment you run, on your own account, any time you doubt a word of this.

6.2 It's been measured, published, and reproduced — this is not one bloke's opinion

- "Lost in the Middle" (Liu et al., 2024), published in a peer-reviewed journal, the Transactions of the Association for Computational Linguistics. The researchers tested six different families of AI model and found the identical pattern in all of them: performance is highest when the needed information sits at the very start or end of the input, and drops sharply — by more than 30% — when that same information is buried in the middle of a long context. Crucially, this held even for models specifically built and sold as "long-context." In plain terms: it isn't that a big context can't be read, it's that the important thing gets lost inside it. https://arxiv.org/abs/2307.03172

- The follow-up benchmarks made it worse, not better. The simplest test of long-context ability is "needle in a haystack" — hide one fact in a big pile and ask the model to find it. Models got good at that, and the hype ran with it. But the RULER benchmark (Hsieh et al., 2024) tested GPT-4, Gemini-1.5 and fifteen open-source models on harder, more realistic tasks and found that passing the easy needle test does not mean the model holds up — almost every model showed large degradation on the harder tasks as the context grew. In fact the researchers found that only about half the models could effectively handle even 32,000 tokens, and almost all fell short of the context length they were advertised to handle. So the reassurance the vendors offer ("look, it handles a million tokens!") is measuring the wrong, easy thing.

- It's common enough to have earned a name. Engineers hit this so routinely in real-world use that in 2025 they coined a term for it: "context rot" — the plain, repeatable observation that these models get less effective as the amount in front of them grows. When a problem gets its own nickname from the people using the tools daily, it has stopped being a theory and become a known hazard.

And the other half of the argument — that Claude re-reads the whole chat each turn and keeps nothing on its own between messages — isn't secret, disputed, or mine. It's simply the published, standard description of how these systems are built: they are stateless, so the entire conversation has to be fed back in every single turn. The saved notes are a separate feature bolted on top, and they too get fed in at the start of every chat. None of this is controversial. It's just not what the prompt-sellers want you looking at, because the moment you understand it, their product stops making sense.

---

7. Being straight about the limits of this

I said I wouldn't overclaim, so here's the honest edge of it.

This document is written in plain English, and plain English rounds off some corners. When I say Claude "re-reads the whole chat" and "grabs the wrong bit," those are faithful pictures of the effect, not a wiring diagram — the real mechanism is attention spread across tokens, described above. When I say it "can't tell what matters from what's done," that's a strong, reliable tendency, not a guarantee it fails every single time — it's "the more you load, the less reliably it copes," not "it always breaks at message ten." And the three-years-and-thousands-of-errors figure is my own experience, offered as testimony, not as a measured statistic.

None of that softens the conclusion. The direction is not in doubt, the effect is published and reproduced, and you can reproduce it yourself. It just means: treat this as a true and useful guide, not a physics equation. Which is exactly what it is.

---

8. What to actually do

1. One job per chat. Do the one thing, get it done, then open a fresh chat for the next. This is the whole method. Everything else is just the reasons it works.

2. Don't leave big things sitting in the chat. Use a long document or file to pull out what you need, then move on — don't let it sit there weighing down every following turn.

3. Keep the saved notes short. A few real facts, never whole documents pasted in — or every chat you open starts weighed down before you begin.

4. Make it prove things. Not sure? It must say so, not dress a guess as fact. Says it's done? It shows you.

5. Never let it guess. If it doesn't know something, it asks or goes and finds out — it does not make it up.

9. If a chat has already gone long and started going wrong

Don't try to rescue it in place. It's already too loaded, and every new message you send just adds to the pile you're trying to escape. Start a fresh chat, and carry over only the few facts the new one genuinely needs: what the job is, where you'd got to, and any decisions already settled. Leave everything else behind. A clean chat carrying the three facts that matter beats a bloated chat carrying everything, every single time.

10. Bottom line

One chat, one task, then start fresh. Keep what's loaded in front of Claude small and it stops making the mess. There is no magic prompt — there was never going to be one, because loading more in is the problem, and a prompt is just more loaded in. There's just this one simple rule, and now you know exactly why it's the only one that works.

---

References

- Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). *Lost in the Middle: How Language Models Use Long Contexts.* Transactions of the Association for Computational Linguistics, 12, 157–173. https://arxiv.org/abs/2307.03172 - Hsieh, C., et al. (2024). *RULER: What's the Real Context Size of Your Long-Context Language Models?* Benchmarked GPT-4, Gemini-1.5 and 15 open-source models; found near-perfect needle-in-a-haystack scores but large degradation on harder tasks as length grows, with only about half handling 32K effectively. https://arxiv.org/abs/2404.06654 - "Cont

Share this post
XFacebookRedditLinkedInWhatsAppTelegramEmail
💬 Got a problem?

aiwebpageseo.com is a data-driven SEO and AEO platform providing a free suite of technical website tools.