r/LocalLLaMA · · 1 min read

FreedomIntelligence/HuatuoGPT-3-27B · Hugging Face

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

FreedomIntelligence/HuatuoGPT-3-27B · Hugging Face

from FreedomIntelligence:

HuatuoGPT-3-27B is a medical LLM built on Qwen3.8-27B with One-stage Policy Optimization (OnePO). OnePO adapts language models to medicine in a single reinforcement-learning stage, without preceding domain-specific supervised fine-tuning. Teacher responses provide temporary guidance and are retired as the model improves.

We release the training code, medical RL dataset, and 8B rubric grader.

(last week they released https://huggingface.co/FreedomIntelligence/HuatuoGPT-3-9B)

submitted by /u/jacek2023
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA