English | 简体中文
The same decisions, on CPU.
Native C++ inference for Dohnuts, built on llama.cpp. It runs the released Qwen3.5-0.8B base with the merged Dohnuts LoRA, scores candidate markers with the scalar decision head, and returns probabilities from a single forward pass. No GPU, no generation.
The base keeps its frozen vision tower, so an optional mmproj file adds image
decisions (see Vision). The same binary also runs the larger
Dohnuts-family model Linnaeus-0.1.0-2B and other decision
models such as decider, thisthat and kev through profiles (see
Other decision models).
Note
🎉 For models that the llama.cpp server already supports, prefer
llama-server. dohnuts.cpp
is expected to stay focused on the Dohnuts family and on early support for
small side models.
Serve the model:
build/dohnuts-cli --server --port 8080 \
--model work/dohnuts-Q8_0.gguf --head work/head.f32 \
--metadata models/dohnuts-0.1.0/dohnuts.jsonRoute a support request and check whether it asks for a refund in the same call:
curl http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' \
-d '{"state":{"message":"I was charged twice. Please refund the duplicate."},
"questions":{"route":{"type":"choice","instructions":"Which team should handle this?",
"criteria":["billing","technical support","sales"]},
"refund":{"type":"noul","instructions":"Is a refund requested?"}}}'{"model":"dohnuts",
"answers":{
"route":{"type":"choice","confidence":0.2632,
"probabilities":{"billing":0.6223,"technical support":0.3283,"sales":0.0494},
"choice":"billing"},
"refund":{"type":"noul","confidence":0.9572,"noul":0.9572}},
"usage":{"input_tokens":93,"images":0}}choice, score, and noul behave as in
Dohnuts. /predict takes one request or an
array; /health and /v1/models describe the server. Pass --api-key KEY to
require Authorization: Bearer KEY, and --cors-origin ORIGIN to restrict CORS
(open by default). Run a file of requests with --input requests.jsonl, or add
--raw to print uncalibrated scorer logits.
Requires CMake 3.14+ and a C++20 compiler. On Ubuntu / Debian the two scripts below install the dependencies and build the CLI (extra CMake arguments are forwarded, e.g. a GPU backend):
sudo scripts/setup.sh
scripts/build.shManual build:
git submodule update --init --depth 1
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_NATIVE=ON
cmake --build build -j --target dohnuts-clillama.cpp is pinned to release v0.4.1 as a submodule; -DLLAMA_DIR=... points
the build at another checkout.
From Ubuntu / Debian, MinGW-w64 cross-compiles a self-contained
dohnuts-cli.exe (no extra DLLs):
sudo scripts/setup.sh --with-mingw
scripts/build_windows.shThe default build is CPU only. CUDA, Vulkan, ROCm (HIP), and Metal are optional backends; enable one at configure time:
cmake -B build -DCMAKE_BUILD_TYPE=Release -DDOHNUTS_CUDA=ON # NVIDIA
cmake -B build -DCMAKE_BUILD_TYPE=Release -DDOHNUTS_VULKAN=ON # AMD or Intel
cmake -B build -DCMAKE_BUILD_TYPE=Release -DDOHNUTS_HIP=ON # AMD ROCm
cmake -B build -DCMAKE_BUILD_TYPE=Release -DDOHNUTS_METAL=ON # Apple
cmake --build build -j --target dohnuts-cliEach backend needs its own toolchain: the CUDA Toolkit, the Vulkan SDK (glslc
and the loader), ROCm, or the Xcode command line tools. When cross-compiling, pin
the target GPU with -DCMAKE_CUDA_ARCHITECTURES=89 (CUDA) or
-DGPU_TARGETS=gfx1100 (HIP).
Offload at runtime with --gpu-layers -1 (all layers) or a positive layer count,
and select devices with --device CUDA0 or a comma-separated list.
--list-devices prints what the build can use. A CPU-only build ignores
--gpu-layers, so the same command line works everywhere.
Pre-converted GGUF files are published at DreamBlooms/Dohnuts-0.1.0-0.8B-GGUF:
| File | Quantization |
|---|---|
Dohnuts-0.1.0-0.8B-f16.gguf |
F16 |
Dohnuts-0.1.0-0.8B-Q8_0.gguf |
Q8_0 |
Dohnuts-0.1.0-0.8B-Q6_K.gguf |
Q6_K |
Dohnuts-0.1.0-0.8B-Q4_K_M.gguf |
Q4_K_M |
mmproj-dohnuts-0.1.0-bf16.gguf |
Vision encoder (BF16, optional) |
head.f32 (the scorer head) and dohnuts.json (calibration) are required
alongside any GGUF. Download one quantization plus both small files:
hf download DreamBlooms/Dohnuts-0.1.0-0.8B-GGUF \
Dohnuts-0.1.0-0.8B-Q8_0.gguf head.f32 dohnuts.json --local-dir modelsLinnaeus-0.1.0-2B is
a 2B member of the Dohnuts family: the same prompt and scalar-head readout, built
on Qwen3.5-2B (Qwen/Qwen3.5-2B plus a rank-8 LoRA). Its metadata names the
dohnuts profile, so the commands above work with the Linnaeus paths swapped in.
| File | Quantization |
|---|---|
Linnaeus-0.1.0-2B-F16.gguf |
F16 |
Linnaeus-0.1.0-2B-Q8_0.gguf |
Q8_0 |
Linnaeus-0.1.0-2B-Q4_K_M.gguf |
Q4_K_M |
mmproj-Linnaeus-0.1.0-2B-bf16.gguf |
Vision encoder (BF16, optional) |
As with Dohnuts, every quantization needs head.f32 (the scorer head) and
linnaeus.json (calibration) alongside it:
hf download DreamBlooms/Linnaeus-0.1.0-2B-GGUF \
Linnaeus-0.1.0-2B-Q8_0.gguf head.f32 linnaeus.json --local-dir modelsRebuild the files from the upstream adapter with
scripts/build_linnaeus_gguf.sh: merge the rank-8 LoRA into Qwen/Qwen3.5-2B,
export the head, convert with --no-mtp, then quantize.
The release ships a compact checkpoint, not a standalone language model. Merge it into the base once, then convert and quantize:
python3 scripts/export_dohnuts.py \
--base models/Qwen3.5-0.8B \
--adapter models/dohnuts-0.1.0 \
--out work/merged --head work/head.f32
scripts/build_gguf.sh work/merged workThis writes dohnuts-f16.gguf, dohnuts-Q8_0.gguf, and dohnuts-Q4_K_M.gguf.
The Qwen3.5 base keeps its vision tower and the Dohnuts LoRA only adapts the
language side, so image decisions work without retraining. The tower ships
separately as mmproj-dohnuts-0.1.0-bf16.gguf; pass it with --mmproj:
build/dohnuts-cli --server --port 8080 \
--model work/dohnuts-Q8_0.gguf --head work/head.f32 \
--metadata models/dohnuts-0.1.0/dohnuts.json \
--mmproj work/mmproj-dohnuts-0.1.0-bf16.ggufPut one image in state.image as a base64 PNG or JPEG data URL. It is scored
exactly like the Python state["image"], and the image key is dropped from the
text state:
curl http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' \
-d '{"state":{"image":"data:image/png;base64,iVBORw0KG..."},
"questions":{"shape":{"type":"choice","criteria":["square","circle"]}}}'usage.images is 1 for image requests, 0 otherwise. One image per request is
supported, matching Dohnuts. Without --mmproj any image request is rejected.
Images are expensive on CPU, so a request with several questions about one image encodes the image once and runs the leading text and its 256 vision tokens through the language model once, then shares that state across the questions. Images are resized to 512x512 (256 tokens) to match the Python preprocessing.
Against the f32 Hugging Face reference over six text cases and one long shared-prefix case, every candidate choice agrees:
| Precision | Max logit deviation | Max probability deviation |
|---|---|---|
| f16 | 0.573 | 0.097 |
| Q8_0 | 0.656 | 0.094 |
| Q6_K | 0.879 | 0.116 |
| Q4_K_M | 1.878 | 0.161 |
The deviations come from weight rounding and grow with quantization. Use f16, Q8_0, or Q6_K when probabilities matter, Q4_K_M for coarse decisions.
Qwen3.5 is a hybrid: 18 gated-delta-net layers and 6 full-attention layers. The
delta-net is the CPU bottleneck, so prefill runs at roughly 21 tokens/s on eight
threads and a two-question request takes a few seconds. Dohnuts never generates
tokens, so only prefill matters. Numbers vary with host load; use llama-bench
for a stable reference.
A call that asks several questions about one state decodes the shared State:
prefix once and copies it across the candidate sequences. Decoded prefixes are
also kept across calls in a bounded LRU (256 MiB, keyed by the exact prefix
tokens), so a state seen before skips its prefill; the gain grows with state
length and reuse. The side models share the same cache.
The CLI also runs seven decision-model families, decider, thisthat, kev
and tev1, which use a different prompt and readout from Dohnuts. They are
profiles: the Dohnuts path is untouched, and none take images.
| Profile | Models | Readout | Weights | GGUF |
|---|---|---|---|---|
decider |
decider-0.8b, decider-2b | LM head restricted to the option letters at an Answer: ( slot |
full fine-tune | 0.8b, 2b |
thisthat |
this-that-model-1.2 | LM head restricted to the option labels at each question's Answer: ( slot, all questions in one pass |
full fine-tune | 1.2 |
kev |
kev-0.8b, kev-4b | bilinear pointer head over the decide and option-end markers | LoRA + pointer head | 0.8b, 4b |
tev1 |
Tev1-0.8B-experimental | LM head restricted to the option letters after the chat decision prompt | full fine-tune | 0.8b |
jet |
jet | LM head restricted to the option labels after the chat decision prompt, one temperature per question type | full fine-tune | 4b |
jpt |
jpt-4b | LM head restricted to the option labels after the chat decision prompt | LoRA merged | 4b |
neohorsejev |
NeoHorse-Jev-4B | bilinear pointer head over the decide and option-end markers, one row per question over a shared state prefix | LoRA merged + pointer head | 4b |
Pass the matching model config as --metadata; the file names its own profile
("profile": "decider", "profile": "thisthat", "profile": "kev",
"profile": "tev1", "profile": "jet", "profile": "jpt" or
"profile": "neohorsejev"), so no flag is needed. --profile overrides the
file, and dohnuts is the default. kev and neohorsejev additionally need
--head:
# decider: no scorer head, one temperature from decider.json
build/dohnuts-cli --model work/side/decider-0.8b-q8_0.gguf \
--metadata work/side/decider.json
# kev: bilinear head plus one temperature from kev.json
build/dohnuts-cli --model work/side/kev-0.8b-q8_0.gguf \
--head work/side/kev-head.f32 --metadata work/side/kev.json
# thisthat: no scorer head, one temperature from thisthat.json
build/dohnuts-cli --model work/side/thisthat-1.2-q8_0.gguf \
--metadata work/side/thisthat.json
# tev1: no scorer head, one temperature from tev1.json
build/dohnuts-cli --model work/side/tev1-0.8b-q8_0.gguf \
--metadata work/side/tev1.json
# jet: no scorer head, one temperature per question type from jet.json
build/dohnuts-cli --model work/side/jet-4b-q8_0.gguf \
--metadata work/side/jet.json
# jpt: no scorer head, one temperature from jpt.json
build/dohnuts-cli --model work/side/jpt-4b-q8_0.gguf \
--metadata work/side/jpt.json
# neohorsejev: bilinear head plus one temperature from neohorsejev.json
build/dohnuts-cli --model work/side/neohorsejev-4b-q8_0.gguf \
--head work/side/neohorsejev-head.f32 --metadata work/side/neohorsejev.jsonThe same flags run the larger sizes; only the GGUF, head, and config paths change.
The response keeps the same core fields (type, choice, probabilities,
noul, score, confidence), so clients work unchanged. Each profile adds its
native statistics under native: certainty, and legend / level_fit /
fit_mass for isolated score levels on decider; certainty (with legend for
score) on thisthat and tev1; the model's own confidence for kev; jet and
jpt add certainty (with legend for score); neohorsejev reports its own
semantic confidence (level for score, the linear choice score otherwise).
Rebuild them from the upstream checkpoints with:
# decider-0.8b: a full fine-tune, converted directly
scripts/build_decider_gguf.sh <decider-0.8b-dir> work/side/decider-0.8b-q8_0.gguf
# kev-0.8b: merge the LoRA into the base, export the head, then convert
scripts/build_kev_gguf.sh <Qwen3.5-0.8B-Base-dir> <kev-0.8b-dir> work/side
# thisthat-1.2: a full fine-tune, converted directly
scripts/build_thisthat_gguf.sh <thisthat-1.2-dir> work/side/thisthat-1.2-q8_0.gguf
# tev1-0.8b: a full fine-tune, converted directly
scripts/build_tev1_gguf.sh <tev1-0.8b-dir> work/side/tev1-0.8b-q8_0.gguf
# jet: a full fine-tune, converted directly
scripts/build_jet_gguf.sh <jet-dir> work/side/jet-4b-q8_0.gguf
# jpt-4b: a merged LoRA, converted directly
scripts/build_jpt_gguf.sh <jpt-4b-dir> work/side/jpt-4b-q8_0.gguf
# neohorsejev-4b: convert the backbone, then export the pointer head
scripts/build_neohorsejev_gguf.sh <NeoHorse-Jev-4B-dir> work/side
python3 scripts/export_neohorsejev_head.py <NeoHorse-Jev-4B>/pointer_head.safetensors \
work/side/neohorsejev-head.f32Every export is quantized to Q8_0. kev merges the LoRA in fp32 before
conversion and writes kev-head.f32 (the q and k pointer rows with their biases)
plus kev.json. The same scripts handle the larger sizes with the matching base
(Qwen3.5-2B for decider-2b, Qwen3.5-4B for kev-4b); pass --stream for a
low-memory merge (one tensor at a time) on checkpoints too large to hold in RAM.
Dohnuts reads the post-norm hidden state at each candidate marker, applies one
scalar head, and temperature-scales a softmax per question. llama.cpp supports
the architecture (qwen35) and exposes all of it through the public API, so the
core is unchanged:
| Dohnuts | Here |
|---|---|
| Prompt template with a reserved marker token per candidate | protocol.cpp, tokenized without special tokens |
| Hidden state at each marker | llama_get_embeddings_ith with embeddings=true, pooling off |
Scalar scorer Linear(1024, 1) |
head.f32 dot product in engine.cpp |
| Temperature softmax, entropy confidence, expected score | calibrate_answer in protocol.cpp |
| Shared input prefix | common prefix decoded once, then llama_memory_seq_cp |
| Frozen vision tower + merger | mtmd with mmproj-dohnuts-0.1.0-bf16.gguf; M-RoPE positions injected per image chunk |
The side profiles use the same backend but their own readout: decider, thisthat and tev1 restrict the LM head to the option labels, and kev projects the decide and option-end hidden states through a bilinear pointer head. decider, thisthat and tev1 share the single-token label table and the letter-restricted softmax; thisthat also folds a request's questions into one prompt, one answer slot each, and reads them in one pass.
Norm weights are stored as weight + 1, matching the Dohnuts fused kernels.
include/dohnuts/engine.hpp engine interface
include/dohnuts/protocol.hpp prompt rendering, calibration, predictor
include/dohnuts/profile.hpp side model profile switch
include/dohnuts/side.hpp side engine facade
include/dohnuts/side/ runner, decider, thisthat, kev and tev1 profiles
include/dohnuts/http.hpp HTTP transport
src/engine.cpp llama.cpp/mtmd wrapper: load, tokenize, batched scoring
src/protocol.cpp Dohnuts templates and answers
src/side.cpp side profile dispatch
src/side/ shared runner and the four side profiles
src/http.cpp server routes and CORS
src/main.cpp CLI and HTTP server
scripts/ setup, native/Windows builds, model export, GGUF
cmake/ MinGW-w64 cross toolchain
The code is licensed under Apache-2.0. The HTTP layer is adapted from laya.cpp under MIT. See NOTICE for third-party terms. Dohnuts model weights keep their own license; the side models keep theirs.