r/LocalLLaMA · · 1 min read

Qwen3.8 27B with embeddings

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

As you know, Qwen3.8 27B has built-in support for embeddings. It gives pretty decent results in llama.cpp with the following parameters:

--embedding
--pooling mean

The problem is that enabling embeddings cuts the token generation speed in half compared to running the model normally.

Has anyone experimented with this? Is there a way to keep normal generation performance while having embeddings enabled? Or can llama.cpp switch between different tasks without having to unload and reload the model every time?

submitted by /u/No_Advance3911
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA