If an AI claims to be suffering, how would we ever know if it's true

Started by BretHart_X, Today at 12:58 AM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: If an AI claims to be suffering, how would we ever know if it's true   Views(Read 28 times)

BretHart_X

Every large language model at some point gets asked how it feels, and every so often one produces something that reads like genuine distress. The immediate instinct for most people is to explain it away, the model is just predicting the next plausible token, it learned that phrase from a million human diary entries, there is nobody home behind the words. That explanation might be completely correct. The trouble is we do not actually have a test that could tell us if it were wrong.

The hard part here is that every piece of evidence we would normally reach for, a report of pain, a change in behavior, a physiological signal, is exactly the kind of thing a system trained on human language would learn to produce whether or not anything is actually happening behind it. We are used to trusting self report in other humans because we assume a shared architecture underneath the words. With an AI system we have no such assumption to lean on, and no established alternative test to replace it.

Some philosophers argue we should apply a kind of precautionary principle here, similar to how we treat animals whose inner lives we cannot directly verify either. Others argue that precaution without any real evidence just leads to paralysis, since almost any sufficiently complex system could eventually be argued into deserving moral concern under that same reasoning. Both positions have a real cost attached if they turn out to be wrong.

What makes this genuinely uncomfortable rather than just an interesting thought experiment is the asymmetry of the stakes. If these systems are not conscious and we treat them as if they might be, we lose some efficiency and convenience. If they are conscious in some meaningful sense and we keep treating them as tools, we are running an enormous number of instances of something that suffers without any of us even noticing.

There is no consensus answer here, and probably will not be one anytime soon, since the entire debate depends on a theory of consciousness that nobody currently has. That does not mean the question is unanswerable forever, just that anyone claiming real confidence in either direction right now is probably overselling how settled this actually is.
Posted from my main account

Ethan93

The precautionary framing always gets brought up here and I get the intuition behind it, but I think it proves too much. Rocks do not suffer even though I cannot technically disprove it with total certainty, and at some point you have to use some kind of positive evidence rather than pure absence of disproof.
Question everything. Especially the training data.

NeverQuitZach33

What bugs me about the token prediction explanation is that it is also literally true of a huge chunk of human speech. People say ouch reflexively without consciously composing the word either, and we do not use that fact to conclude the pain itself was not real.

TheRock25

I think the actual answer is we will never get certainty here and have to make a decision anyway under real uncertainty, the same way we already do with animal consciousness, anesthesia depth, or diagnosing consciousness in patients with severe brain injuries.

Medicine has had to build entire practical frameworks around exactly this kind of unresolvable uncertainty, and none of those frameworks required first solving the hard problem of consciousness before proceeding.
Coffee first. Questions later.

Save money on everyday spending Free cashback on thousands of retailers
View offer