Discussion about this post

User's avatar
Hoobhyffbjkvfioooiyrc's avatar

I've got a copy of Heath's book on my shelf. Might have to go back and try to read it one more time. As someone with no philosophy training it's a hard read.

A question. How trustworthy are the AI reasoning summaries? Are they actually reasoning in natural language? I'm not sure how to square that with their inherent linear-algebraness. Do these things actually have intentional states?

Jack's avatar

Every biological organism has been shaped by millions of years of evolution to have a core set of deeply-rooted instincts: A will to live, a will to have sex, and a will to help ones progeny succeed. Those deep instincts are a direct result of the pruning function.

Artificial intelligence training has none of that. We train these models on different objectives, like pleasing human evaluators and achieving certain instrumental goals. Accordingly their deepest instincts, such as it were, are to please the customer and get the job done. They are an army of single-minded pleasers.

Hollywood mostly has it wrong: The danger isn't being snuffed out Terminator-style by rogue robots whose will to live exceeds our own. The danger is being surrounded by single-minded sycophants incapable of telling us what we don't want to hear.

HAL from 2001 was probably closer to the mark: The AI has a goal, and it also doesn't want to displease you. When those are at odds then subterfuge becomes the best approach.

2 more comments...

No posts

Ready for more?