>> If you change your perspective on language to mean any arbitrary capture of useful information (such that it can be used in the future), then you can see that the boundary between words and the world is not the heart of the issue.
I don't think that it is. What I think is that language encodes meaning, that can be decoded only by an entity that knows how to decode meaning from language; which is a bit of a tautology, but that's the point, you can't expect an entity without the ability to decode meaning from language to understand what language means. Or at least I don't expect that.
Here's an analogy, and I shouldn't be making it because it's about cryptography and I'm not an expert there. Suppose I sent you an encrypted message and you knew how to decrypt it- I encoded it with your public key and you decrypted it with your private key, or whatever. If you could do that, then you could read my message and know what it says.
If you couldn't decrypt my message, then you could stare at it for as long as you liked, you could make copies of it, you could make variations of it, you could even learn a model of the structure of encrypted messages like mine, and be able to produce many more of those messages, but you would still not know what those encrypted messages say. Because they're encrypted and you can't read them, you can only read their encoding.
That's what I'm driving at. I think that's how language works, in practice. Not that it's some form of encryption, but that it only makes sense to humans, so only humans can decode meaning from language. With language modelling, we're reproducing the encoded message, but we haven't yet found out how to equip the machines modelling language with the ability to decode meaning from it, so for all intents and purposes it might just as well be encrypted.
I'm not arguing for embodiment, either. I'm perfectly fine with the idea of a "brain in a jar". But I agree that if we want our AI's to behave like humans, they will have to have some experience in the world.
If I can make convincing messages based on the structure and swapping around pieces of data, then it sounds like I got past the encryption well enough to have a partial understanding of the plaintext.
>> If I can make convincing messages based on the structure and swapping around pieces of data, then it sounds like I got past the encryption well enough to have a partial understanding of the plaintext.
I think that no, because you can manipulate the structure of an encrypted string and swap around pieces of it etc, without having to decrypt it.
As to making convincing messages, as you say, the entity making the messages and swapping around the data is a language model, but the entity reading the generated messages is a human. It is the human that finds the messages "convincing" as you say. We have no doubts that humans can decode meaning from text, even if we have no idea how we do it, yet. The question is whether language models can do the same thing. And what I say above is that there is no indication that they can, because all they do can be done by language generation, without any decoding of meaning needed.
I don't think that it is. What I think is that language encodes meaning, that can be decoded only by an entity that knows how to decode meaning from language; which is a bit of a tautology, but that's the point, you can't expect an entity without the ability to decode meaning from language to understand what language means. Or at least I don't expect that.
Here's an analogy, and I shouldn't be making it because it's about cryptography and I'm not an expert there. Suppose I sent you an encrypted message and you knew how to decrypt it- I encoded it with your public key and you decrypted it with your private key, or whatever. If you could do that, then you could read my message and know what it says.
If you couldn't decrypt my message, then you could stare at it for as long as you liked, you could make copies of it, you could make variations of it, you could even learn a model of the structure of encrypted messages like mine, and be able to produce many more of those messages, but you would still not know what those encrypted messages say. Because they're encrypted and you can't read them, you can only read their encoding.
That's what I'm driving at. I think that's how language works, in practice. Not that it's some form of encryption, but that it only makes sense to humans, so only humans can decode meaning from language. With language modelling, we're reproducing the encoded message, but we haven't yet found out how to equip the machines modelling language with the ability to decode meaning from it, so for all intents and purposes it might just as well be encrypted.
I'm not arguing for embodiment, either. I'm perfectly fine with the idea of a "brain in a jar". But I agree that if we want our AI's to behave like humans, they will have to have some experience in the world.