Which local model is actually good at knowing when to stop and ask you a question?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
I’ve been thinking about this after using more agentic/local coding models.
A lot of the newer models are surprisingly good at continuing on their own.
But sometimes that seems like the problem.
If a requirement is ambiguous, I’d rather the model stop and ask:
“Do you mean A or B?”
instead of spending 10 minutes reasoning, making an assumption, calling tools and then confidently building the wrong thing.
I don’t see this behavior discussed much in benchmarks either. We measure coding, reasoning, tool use, context length, etc., but not really whether a model knows when it doesn’t have enough information to continue.
My genuine question for people running models locally every day is,
Which model have you found best at this?
And is it mostly the model itself, the system prompt, or your agent harness that makes the difference?
[link] [comments]
More from r/LocalLLaMA
-
NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not.
Sep 28
-
3090 for $1500???
Sep 28
-
modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection
Sep 28
-
Minisforum MS-S1 MAX-P495 @ €7.799,00
Sep 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.