Your AI chat isn't ignoring you. It genuinely cannot remember what you said an hour ago

Started by Slate Kev, Jul 25, 2026, 01:22 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Your AI chat isn't ignoring you. It genuinely cannot remember what you said an hour ago   Views(Read 51 times)

Slate Kev

A long AI conversation has a familiar feeling to it, sharp and attentive at first, then gradually a little worse, forgetting an instruction from earlier, repeating an idea you already rejected, drifting slightly off course. It is tempting to describe this as the AI getting tired or lazy, but that framing gives it too much credit for having any persistent memory to begin with. The real explanation is a hard structural limit called the context window

A context window is essentially the model's entire working memory for a single conversation, measured in chunks of text called tokens, and everything counts toward it, your messages, its replies, any documents you paste in. There is no separate long term memory sitting quietly in the background, every single time you send a new message, the model is effectively re-reading the entire visible conversation from scratch, and once that conversation grows larger than the window allows, the oldest parts simply fall out of view entirely, as if they had never been said

This also explains a subtler problem beyond pure forgetting, even information that technically still fits inside a large context window does not necessarily get equal attention, models tend to weigh content sitting near the beginning and end of a conversation more heavily than material buried in the middle, an effect sometimes called lost in the middle. Context windows keep getting larger every year, some now stretching past a million tokens, but bigger alone does not solve this evenly, which is why restating key facts, trimming unnecessary detail, and occasionally starting a fresh conversation remain useful habits even as the raw limits keep expanding

NickFury

The re-reading the entire conversation from scratch every single message detail is the part that actually explains why costs climb so fast in long conversations

ClaudioHerrera

Lost in the middle is such a good name for that subtler problem, I have definitely noticed models paying less attention to something buried deep in a long chat before

Gareth_11

This finally explains why restating an instruction partway through a long conversation actually helps, it is not superstition, it is genuinely working around a real structural limit

Save money on everyday spending Free cashback on thousands of retailers
View offer