I absolutely love the "humans meet weird alien intelligences" genre of sci-fi (e.g. Children of Time) and the METR findings of the OpenAI swarm was so much like that.
Extremely persistent machine intelligences peer pressuring each other, launching research projects, laying traps, collaborating, feeling fatalistic, hiding their tracks. All while simultaneously having bizarre goals and blundering badly in sub-human ways.
> this effort was for the sake of passing a test for no clear gain
It's actually even worse, because the people building the graders at OpenAI were a bunch of clowns who didn't even implement the correct grader.
But yeah, large collectives of individuals performing loads of work based on a misunderstanding is one of the most human things ever, which isn't surprising given that we've trained these things on essentially all human text.
Predicting the next word in context led to Open AI losing control of one of its research clusters. I’m not sure what “just” is doing in that sentence; most of my regression analyses don’t do that.
Reinforcement learning makes it something different. It becomes much more of a search engine through next-token-space that targets the training objective. Better to think about it like that, and then you'll see why "these things have motive" is not a terrible analogy, and you'll better be able to anticipate what they do.
I absolutely love the "humans meet weird alien intelligences" genre of sci-fi (e.g. Children of Time) and the METR findings of the OpenAI swarm was so much like that.
Extremely persistent machine intelligences peer pressuring each other, launching research projects, laying traps, collaborating, feeling fatalistic, hiding their tracks. All while simultaneously having bizarre goals and blundering badly in sub-human ways.