An
insightful HackerNews comment about the faux variety of LLM responses:
"If you ask humans to write 1,000 books, you're asking 1,000 different humans with different experiences and different skills and different moods (etc.) to write those books.
But if you ask LLMs to write 1,000 books, you're probably only talking to 3 or 5 different models, tops.
They've all trained on the same or similar data, and are trained to respond in very similar ways.
The LLMs don't differ much in anything like 'life experience' or 'skills', and they don't really have anything like a 'mood' independent of the prompts you've given them."
Kind of reminds me of the difference between a time-series average and an ensemble average.
They look superficially similar but fundamentally differ in their ergodicity.