> this effort was for the sake of passing a test for no clear gain
It's actually even worse, because the people building the graders at OpenAI were a bunch of clowns who didn't even implement the correct grader.
But yeah, large collectives of individuals performing loads of work based on a misunderstanding is one of the most human things ever, which isn't surprising given that we've trained these things on essentially all human text.
The only thing there that's weird from a human perspective is that all this effort was for the sake of passing a test for no clear gain.