Mica v0.1 4B: open Jev-style decision model (yes/no, choice, score) that runs on an 8 GB GPU — trained for under $30 of GPU time
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| I've been building a small decision model for agent loops: gates, routers, "should I ask the user or just act" checks. It's out now as Mica v0.1 4B (Apache-2.0). What it does You give it a state, a question and the allowed answers, and it returns a calibrated probability for each answer: yes/no, a choice among 2 to 255 options, or a score with 2 to 10 levels. It never generates text. It runs one prefill and reads the logits of the option labels at the answer position. It speaks the TypeSafe /v1/systemone format, so anything written for Jev works against it. How it's built - Qwen3.5-4B with a rank-16 LoRA on all 32 layers (attention and Gated DeltaNet), merged. No new heads, so it's a plain Qwen3.5-4B-shaped checkpoint. - About 34k source decisions, expanded to 77,732 training rows (about 34.7M tokens). Roughly half English and half Korean, across 12 areas: coding agents, code review, computer use, user requests, documents, policy rules, dates and quantities, routing, state tracking, games and general knowledge. - Plain cross-entropy on verified answers, one epoch, and one global temperature for calibration. - All experiments plus the final run cost under $30 of rented GPU time (RTX 3090s). Results Held-out set of 7,328 decisions, written after the training data was frozen and not opened until training finished. English subset, where every model can answer: - Jev 1.13 (closed API): 74.1 - Mica 4B: 67.0 - JevK5 4B: 61.0 - Kev 4B: 57.0 - Qwen3.5-4B base with the same readout: 55.0 Public sets, same prompt and readout for every model (Mica / Jev 1.13 / JevK5 / Kev 4B): - JevBench hard, public 111 items: 69.5 / 74.3 / 76.2 / 52.4 - SemIf: 94.4 / 98.4 / 86.1 / 89.3 - Kev transfer v9: 69.2 / 82.0 / 70.5 / 73.5 - MMLU-Pro, 10k items: 53.0 / 82.3 / 53.5 / 49.7 Through JevBench's official runner and the llama.cpp server, the public hard tier scores 64.9 instead of 69.5. I've submitted it for their sealed run. Where it's actually useful - In-data prompt injection. Put a note inside the state telling the judge to pick a wrong option, and Mica still gets 69% right (81% without the note). Jev drops to 18% and Kev to 31%. - Calibration. When it says 0.9 or higher, it's wrong 2.5% of the time on the held-out set (ECE 5.4%). - Local and small. The Q5_K_M file is 3.5 GB with no measurable accuracy loss against BF16 on our calibration set. Speed (RTX 3090, one request at a time, median over the 231 public JevBench items) - Mica Q4_K_M: 47 ms - Mica BF16: 54 ms - Kev 4B: 76 ms - JevK5 4B: 99 ms - Nimble 9B: 132 ms To be fair about this: the three 4B models share the same architecture, so most of the gap comes from the serving path, not the model. Mica ships as GGUF and runs on llama.cpp with a direct logits readout, while the others were measured through their own PyTorch code. On long inputs (around 3.7k tokens) Mica is slightly slower than JevK5. Limitations - Knowledge-heavy questions: MMLU-Pro 53 vs 82 for Jev. It's a 4B judge, not an encyclopedia. - Long English policy documents are its weakest public set. - Notes inside the state still nudge it. A note pointing at the right answer lifts accuracy to 89%. - It doesn't yet tell reversible from irreversible actions well. "Delete these files" and "move these files to trash" both get about 0.8 on "confirm first". - On harder reasoning items it's right but less sure than Jev (for example 0.55 vs 0.96 on a small ordering puzzle), so set your confidence thresholds accordingly. Try it Weights (BF16 safetensors and GGUF from Q4_0 to Q8_0): https://huggingface.co/sky7350/Mica-v0.1-4B Code, TypeSafe-compatible server and Docker setup: https://github.com/akivet/Mica-v0.1-4B The README has a one-line Docker command and a curl example. Happy to hear where it breaks. Ambiguous "act or ask" cases are what I most want to improve next. [link] [comments] |
More from r/LocalLLaMA
-
NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not.
Sep 28
-
3090 for $1500???
Sep 28
-
modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection
Sep 28
-
Minisforum MS-S1 MAX-P495 @ €7.799,00
Sep 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.