r/LocalLLaMA · · 1 min read

Why are AI model tests always the same generic prompts?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Okay, hear me out. Why is it that every time a new model comes out, all the tests I see are "make a car game," "make a website," or something equally generic, usually from a prompt that's barely a line and a half long?

That doesn't feel like a fair test. I'd be way more interested in seeing evaluations with detailed, real world instructions, the kind of complex tasks you'd actually run into on the job.

From what I've looked into, most benchmarks rely on simple multiple choice or short coding problems that are easy to auto-score. The only one that seems to get close to real-world work is deepswe which looks okeyish. Everything else feels pretty shallow. Did I miss some?

And even youtubers, most of them just run the same lazy one line prompts and spend half of the time screaming at the screen..

submitted by /u/ddeeppiixx
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA