Rendered at 02:15:13 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
nl 2 hours ago [-]
> Epistemological weirdness
> They are also notably bad at judging the historical significance of what they find.
I use LLMs for some things that are outside the more common use-cases (in my case 3D design for 3D printing) and one thing I've noticed is that the errors it makes are so completely unlike human errors that they are hard to anticipate.
It will do things like build perfect snap catches but put them so the the pieces they are connecting are rotated 90 degrees from how they should be. It's "dumb" error, but hard to say the model itself if dumb because it does other very hard things so perfectly.
> seven chord groups
This sounds a lot more like Opus 5.0 than Opus 5.5 TBH. I wonder if that was an earlier investigation because 5.5 has improved that kind of language a lot.
Waterluvian 1 hours ago [-]
This analogy may be too close to the real thing to work, but it reminds me of a Chinese room type situation where its entire understanding of the world is through messages of text.
You say that’s an error a human couldn’t do, but imagine if the human has never seen or touched the kind of item you were making and relied entirely on text descriptions to build its ontology. Off by 90 seems like such a believable mistake.
meowface 45 minutes ago [-]
Also, not to sound like a naive hypemonger, but: in a decade I'd bet a ton of money the best AI systems will make strange mistakes of this nature at a far, far lower rate than they do today. They will gain a more holistic and more human-like perspective about each task.
(even if it's through some silly means like explicitly talking to themselves like "if I were a human doing this, what [... 5 million tokens in 2 seconds ...]" but also of course if they crack ASI and get something more efficient and intelligent than a human brain by then)
xmprt 23 minutes ago [-]
I think of it kind of like how Chess AI make "mistakes" which are unrecognizable to humans but a stronger AI would be able to pick them apart. That's kind of scary...
BoppreH 24 minutes ago [-]
AI capabilities are "spiky": they extend far in some dimensions but fall short in others, seemingly at random. See for example the recent "thus spoke compute" musical[1]. It's an absolute banger, the graphics are impressive, and so is the writing. But some of the metaphors make no sense, the text highlights are in the wrong places, and the train animation at 2:35 is running backwards!
A person capable of making the rest of the video would never make those mistakes, but an AI does. Perhaps our intelligence is also spiky, and we're just used to the general shape and variance within humans.
It's "dumb" error, but hard to say the model itself if dumb because it does other very hard things so perfectly.
Maybe the model isn't intelligence in any form, except perhaps as an imperfect reflection of the intelligence of its training data.
Isamu 22 minutes ago [-]
I agree, there’s the collective intelligence that created all the content used to train the model. The model is a superposition of all that material with RL tuning. Analogously to reading a book, the intelligence you perceive is from the book’s creator.
morpheos137 1 hours ago [-]
In general llms are weak with spatial reasoning. This seems to be an unsolved problem. Probably because human language is generally imprecise spatially and humans think about spatial problems in visual terms. I wonder if having an llm make a 3d design in a format an image model could check would result in a better outcome?
Aerroon 17 minutes ago [-]
Is it that LLMs are weak with spatial reasoning (and memory) or is it that we are unusually good at it?
When I need to use a program I seldomly use I'm far more likely to remember where I need to click to open it than the word I need to search for to open it.
NiloCK 29 minutes ago [-]
I think that this is an unsolved problem in the same way that mangled fingers in image generation was an unsolved problem.
Through at least Opus 4, LLMs were practically useless for authoring any sort of coherent procedural closed-curve geometry (I know this with strong confidence because of the little animated guys at https://letterspractice.com).
Opus 5.5 can bang it all out. Possibly a deliberate RL sort of thing or maybe another surprise emergent capability.
early_exit 1 hours ago [-]
Good read! How many pages of text were in scope? I'm not sure if the 1615+1629 pages were the total or just a subagent.
If they were the total I would say it was arguably more impressive the author was able to narrow it down to just 3000 pages than it was to find the dodo mention amongst those!
mkl 31 minutes ago [-]
Those are years, not page counts, right? The article mentions "millions of records", but I'm not sure how big a record can be.
dgellow 3 hours ago [-]
What a great read, I generally associate substack with verbose, low quality content, but definitely not the case here!
> They are also notably bad at judging the historical significance of what they find.
I use LLMs for some things that are outside the more common use-cases (in my case 3D design for 3D printing) and one thing I've noticed is that the errors it makes are so completely unlike human errors that they are hard to anticipate.
It will do things like build perfect snap catches but put them so the the pieces they are connecting are rotated 90 degrees from how they should be. It's "dumb" error, but hard to say the model itself if dumb because it does other very hard things so perfectly.
> seven chord groups
This sounds a lot more like Opus 5.0 than Opus 5.5 TBH. I wonder if that was an earlier investigation because 5.5 has improved that kind of language a lot.
You say that’s an error a human couldn’t do, but imagine if the human has never seen or touched the kind of item you were making and relied entirely on text descriptions to build its ontology. Off by 90 seems like such a believable mistake.
(even if it's through some silly means like explicitly talking to themselves like "if I were a human doing this, what [... 5 million tokens in 2 seconds ...]" but also of course if they crack ASI and get something more efficient and intelligent than a human brain by then)
A person capable of making the rest of the video would never make those mistakes, but an AI does. Perhaps our intelligence is also spiky, and we're just used to the general shape and variance within humans.
[1] https://www.youtube.com/watch?v=Cq8qO-NjYIg
Maybe the model isn't intelligence in any form, except perhaps as an imperfect reflection of the intelligence of its training data.
When I need to use a program I seldomly use I'm far more likely to remember where I need to click to open it than the word I need to search for to open it.
Through at least Opus 4, LLMs were practically useless for authoring any sort of coherent procedural closed-curve geometry (I know this with strong confidence because of the little animated guys at https://letterspractice.com).
Opus 5.5 can bang it all out. Possibly a deliberate RL sort of thing or maybe another surprise emergent capability.
If they were the total I would say it was arguably more impressive the author was able to narrow it down to just 3000 pages than it was to find the dodo mention amongst those!