r/LocalLLaMA · · 2 min read

To my surprise I found gemma4 much better at tool-calling than Qwen

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

I've had fairly good luck getting off the ground coding at home with both qwen3.6 35B a3b, and also qwen3.8 27B. However once I switched from a chat window (where the robot wrote code blocks that I could copy/paste into a text editor) to a simple agentic loop, things got funny. My agentic program instructs the robot what utilities they have available, and after their chat they can include a solitary line with just ---, and then a list of commands they'd like to run on my computer, prepended with "RUN: ". I can then approve or edit the commands, and then the robot gets the results, and this repeats until the robot proposes that the goal is complete, and I sign off on it or provide a comment to steer it in the right direction.

What has happened in this experience is surprising me. My go-to models, Qwen3.6 and Qwen3.8, are utterly unable to follow the instructions consistently, and get themselves in insane feedback loops once they clog up their context. Like instead of my suggestion, why don't you `ast-grep outline --lang CSS foobar.html` they will just cat the whole file, and then do it again on the next interaction, and then they started interspersing raw thinking tokens into the output and even some of the Qwen's template delimiters started leaking into the output.

Then, I did something I thought I would never do: pulled gemma4:12b out of the trash and found gemma4 very obedient and quick to take up the instructions and follow them accurately. That led me to download gemma4:26b and it does the same, although it's a little better at planning its work and proposing ways forward.

I even went all the way back to Qwen2.5:3b-instruct as ChatGPT told me ages ago, this model is trained especially for tool-calling.

So, maybe I'm living in the stone ages but I like to treat all this AI stuff as a learning experience and that's why I prefer to grind for weeks with DIY projects instead of downloading something. My implicit question is: Can we discuss tool-calling models that you thought worked better or worse for you? I'm wondering if there's a trade-off slider between "creative" and "obedient" when it comes to following rules vs. finding innovative solutions.

submitted by /u/spammmmmmmmy
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA