Skip to content

Repository files navigation

dohnuts.cpp

English | 简体中文

The same decisions, on CPU.

Native C++ inference for Dohnuts, built on llama.cpp. It runs the released Qwen3.5-0.8B base with the merged Dohnuts LoRA, scores candidate markers with the scalar decision head, and returns probabilities from a single forward pass. No GPU, no generation.

The base keeps its frozen vision tower, so an optional mmproj file adds image decisions (see Vision). The same binary also runs the larger Dohnuts-family model Linnaeus-0.1.0-2B and other decision models such as decider, thisthat and kev through profiles (see Other decision models).

Note

🎉 For models that the llama.cpp server already supports, prefer llama-server. dohnuts.cpp is expected to stay focused on the Dohnuts family and on early support for small side models.

One message, several decisions

Serve the model:

build/dohnuts-cli --server --port 8080 \
  --model work/dohnuts-Q8_0.gguf --head work/head.f32 \
  --metadata models/dohnuts-0.1.0/dohnuts.json

Route a support request and check whether it asks for a refund in the same call:

curl http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' \
  -d '{"state":{"message":"I was charged twice. Please refund the duplicate."},
       "questions":{"route":{"type":"choice","instructions":"Which team should handle this?",
         "criteria":["billing","technical support","sales"]},
         "refund":{"type":"noul","instructions":"Is a refund requested?"}}}'
{"model":"dohnuts",
 "answers":{
   "route":{"type":"choice","confidence":0.2632,
     "probabilities":{"billing":0.6223,"technical support":0.3283,"sales":0.0494},
     "choice":"billing"},
   "refund":{"type":"noul","confidence":0.9572,"noul":0.9572}},
 "usage":{"input_tokens":93,"images":0}}

choice, score, and noul behave as in Dohnuts. /predict takes one request or an array; /health and /v1/models describe the server. Pass --api-key KEY to require Authorization: Bearer KEY, and --cors-origin ORIGIN to restrict CORS (open by default). Run a file of requests with --input requests.jsonl, or add --raw to print uncalibrated scorer logits.

Build

Requires CMake 3.14+ and a C++20 compiler. On Ubuntu / Debian the two scripts below install the dependencies and build the CLI (extra CMake arguments are forwarded, e.g. a GPU backend):

sudo scripts/setup.sh
scripts/build.sh

Manual build:

git submodule update --init --depth 1
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_NATIVE=ON
cmake --build build -j --target dohnuts-cli

llama.cpp is pinned to release v0.4.1 as a submodule; -DLLAMA_DIR=... points the build at another checkout.

Windows (cross-compile)

From Ubuntu / Debian, MinGW-w64 cross-compiles a self-contained dohnuts-cli.exe (no extra DLLs):

sudo scripts/setup.sh --with-mingw
scripts/build_windows.sh

GPU

The default build is CPU only. CUDA, Vulkan, ROCm (HIP), and Metal are optional backends; enable one at configure time:

cmake -B build -DCMAKE_BUILD_TYPE=Release -DDOHNUTS_CUDA=ON     # NVIDIA
cmake -B build -DCMAKE_BUILD_TYPE=Release -DDOHNUTS_VULKAN=ON   # AMD or Intel
cmake -B build -DCMAKE_BUILD_TYPE=Release -DDOHNUTS_HIP=ON      # AMD ROCm
cmake -B build -DCMAKE_BUILD_TYPE=Release -DDOHNUTS_METAL=ON    # Apple
cmake --build build -j --target dohnuts-cli

Each backend needs its own toolchain: the CUDA Toolkit, the Vulkan SDK (glslc and the loader), ROCm, or the Xcode command line tools. When cross-compiling, pin the target GPU with -DCMAKE_CUDA_ARCHITECTURES=89 (CUDA) or -DGPU_TARGETS=gfx1100 (HIP).

Offload at runtime with --gpu-layers -1 (all layers) or a positive layer count, and select devices with --device CUDA0 or a comma-separated list. --list-devices prints what the build can use. A CPU-only build ignores --gpu-layers, so the same command line works everywhere.

Models

Pre-converted GGUF files are published at DreamBlooms/Dohnuts-0.1.0-0.8B-GGUF:

File Quantization
Dohnuts-0.1.0-0.8B-f16.gguf F16
Dohnuts-0.1.0-0.8B-Q8_0.gguf Q8_0
Dohnuts-0.1.0-0.8B-Q6_K.gguf Q6_K
Dohnuts-0.1.0-0.8B-Q4_K_M.gguf Q4_K_M
mmproj-dohnuts-0.1.0-bf16.gguf Vision encoder (BF16, optional)

head.f32 (the scorer head) and dohnuts.json (calibration) are required alongside any GGUF. Download one quantization plus both small files:

hf download DreamBlooms/Dohnuts-0.1.0-0.8B-GGUF \
  Dohnuts-0.1.0-0.8B-Q8_0.gguf head.f32 dohnuts.json --local-dir models

Linnaeus-0.1.0-2B

Linnaeus-0.1.0-2B is a 2B member of the Dohnuts family: the same prompt and scalar-head readout, built on Qwen3.5-2B (Qwen/Qwen3.5-2B plus a rank-8 LoRA). Its metadata names the dohnuts profile, so the commands above work with the Linnaeus paths swapped in.

File Quantization
Linnaeus-0.1.0-2B-F16.gguf F16
Linnaeus-0.1.0-2B-Q8_0.gguf Q8_0
Linnaeus-0.1.0-2B-Q4_K_M.gguf Q4_K_M
mmproj-Linnaeus-0.1.0-2B-bf16.gguf Vision encoder (BF16, optional)

As with Dohnuts, every quantization needs head.f32 (the scorer head) and linnaeus.json (calibration) alongside it:

hf download DreamBlooms/Linnaeus-0.1.0-2B-GGUF \
  Linnaeus-0.1.0-2B-Q8_0.gguf head.f32 linnaeus.json --local-dir models

Rebuild the files from the upstream adapter with scripts/build_linnaeus_gguf.sh: merge the rank-8 LoRA into Qwen/Qwen3.5-2B, export the head, convert with --no-mtp, then quantize.

Build the GGUF yourself

The release ships a compact checkpoint, not a standalone language model. Merge it into the base once, then convert and quantize:

python3 scripts/export_dohnuts.py \
    --base models/Qwen3.5-0.8B \
    --adapter models/dohnuts-0.1.0 \
    --out work/merged --head work/head.f32

scripts/build_gguf.sh work/merged work

This writes dohnuts-f16.gguf, dohnuts-Q8_0.gguf, and dohnuts-Q4_K_M.gguf.

Vision

The Qwen3.5 base keeps its vision tower and the Dohnuts LoRA only adapts the language side, so image decisions work without retraining. The tower ships separately as mmproj-dohnuts-0.1.0-bf16.gguf; pass it with --mmproj:

build/dohnuts-cli --server --port 8080 \
  --model work/dohnuts-Q8_0.gguf --head work/head.f32 \
  --metadata models/dohnuts-0.1.0/dohnuts.json \
  --mmproj work/mmproj-dohnuts-0.1.0-bf16.gguf

Put one image in state.image as a base64 PNG or JPEG data URL. It is scored exactly like the Python state["image"], and the image key is dropped from the text state:

curl http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' \
  -d '{"state":{"image":"data:image/png;base64,iVBORw0KG..."},
       "questions":{"shape":{"type":"choice","criteria":["square","circle"]}}}'

usage.images is 1 for image requests, 0 otherwise. One image per request is supported, matching Dohnuts. Without --mmproj any image request is rejected.

Images are expensive on CPU, so a request with several questions about one image encodes the image once and runs the leading text and its 256 vision tokens through the language model once, then shares that state across the questions. Images are resized to 512x512 (256 tokens) to match the Python preprocessing.

Accuracy

Against the f32 Hugging Face reference over six text cases and one long shared-prefix case, every candidate choice agrees:

Precision Max logit deviation Max probability deviation
f16 0.573 0.097
Q8_0 0.656 0.094
Q6_K 0.879 0.116
Q4_K_M 1.878 0.161

The deviations come from weight rounding and grow with quantization. Use f16, Q8_0, or Q6_K when probabilities matter, Q4_K_M for coarse decisions.

Speed

Qwen3.5 is a hybrid: 18 gated-delta-net layers and 6 full-attention layers. The delta-net is the CPU bottleneck, so prefill runs at roughly 21 tokens/s on eight threads and a two-question request takes a few seconds. Dohnuts never generates tokens, so only prefill matters. Numbers vary with host load; use llama-bench for a stable reference.

A call that asks several questions about one state decodes the shared State: prefix once and copies it across the candidate sequences. Decoded prefixes are also kept across calls in a bounded LRU (256 MiB, keyed by the exact prefix tokens), so a state seen before skips its prefill; the gain grows with state length and reuse. The side models share the same cache.

Other decision models

The CLI also runs seven decision-model families, decider, thisthat, kev and tev1, which use a different prompt and readout from Dohnuts. They are profiles: the Dohnuts path is untouched, and none take images.

Profile Models Readout Weights GGUF
decider decider-0.8b, decider-2b LM head restricted to the option letters at an Answer: ( slot full fine-tune 0.8b, 2b
thisthat this-that-model-1.2 LM head restricted to the option labels at each question's Answer: ( slot, all questions in one pass full fine-tune 1.2
kev kev-0.8b, kev-4b bilinear pointer head over the decide and option-end markers LoRA + pointer head 0.8b, 4b
tev1 Tev1-0.8B-experimental LM head restricted to the option letters after the chat decision prompt full fine-tune 0.8b
jet jet LM head restricted to the option labels after the chat decision prompt, one temperature per question type full fine-tune 4b
jpt jpt-4b LM head restricted to the option labels after the chat decision prompt LoRA merged 4b
neohorsejev NeoHorse-Jev-4B bilinear pointer head over the decide and option-end markers, one row per question over a shared state prefix LoRA merged + pointer head 4b

Pass the matching model config as --metadata; the file names its own profile ("profile": "decider", "profile": "thisthat", "profile": "kev", "profile": "tev1", "profile": "jet", "profile": "jpt" or "profile": "neohorsejev"), so no flag is needed. --profile overrides the file, and dohnuts is the default. kev and neohorsejev additionally need --head:

# decider: no scorer head, one temperature from decider.json
build/dohnuts-cli --model work/side/decider-0.8b-q8_0.gguf \
  --metadata work/side/decider.json

# kev: bilinear head plus one temperature from kev.json
build/dohnuts-cli --model work/side/kev-0.8b-q8_0.gguf \
  --head work/side/kev-head.f32 --metadata work/side/kev.json

# thisthat: no scorer head, one temperature from thisthat.json
build/dohnuts-cli --model work/side/thisthat-1.2-q8_0.gguf \
  --metadata work/side/thisthat.json

# tev1: no scorer head, one temperature from tev1.json
build/dohnuts-cli --model work/side/tev1-0.8b-q8_0.gguf \
  --metadata work/side/tev1.json

# jet: no scorer head, one temperature per question type from jet.json
build/dohnuts-cli --model work/side/jet-4b-q8_0.gguf \
  --metadata work/side/jet.json

# jpt: no scorer head, one temperature from jpt.json
build/dohnuts-cli --model work/side/jpt-4b-q8_0.gguf \
  --metadata work/side/jpt.json

# neohorsejev: bilinear head plus one temperature from neohorsejev.json
build/dohnuts-cli --model work/side/neohorsejev-4b-q8_0.gguf \
  --head work/side/neohorsejev-head.f32 --metadata work/side/neohorsejev.json

The same flags run the larger sizes; only the GGUF, head, and config paths change.

The response keeps the same core fields (type, choice, probabilities, noul, score, confidence), so clients work unchanged. Each profile adds its native statistics under native: certainty, and legend / level_fit / fit_mass for isolated score levels on decider; certainty (with legend for score) on thisthat and tev1; the model's own confidence for kev; jet and jpt add certainty (with legend for score); neohorsejev reports its own semantic confidence (level for score, the linear choice score otherwise).

Rebuild them from the upstream checkpoints with:

# decider-0.8b: a full fine-tune, converted directly
scripts/build_decider_gguf.sh <decider-0.8b-dir> work/side/decider-0.8b-q8_0.gguf

# kev-0.8b: merge the LoRA into the base, export the head, then convert
scripts/build_kev_gguf.sh <Qwen3.5-0.8B-Base-dir> <kev-0.8b-dir> work/side

# thisthat-1.2: a full fine-tune, converted directly
scripts/build_thisthat_gguf.sh <thisthat-1.2-dir> work/side/thisthat-1.2-q8_0.gguf

# tev1-0.8b: a full fine-tune, converted directly
scripts/build_tev1_gguf.sh <tev1-0.8b-dir> work/side/tev1-0.8b-q8_0.gguf

# jet: a full fine-tune, converted directly
scripts/build_jet_gguf.sh <jet-dir> work/side/jet-4b-q8_0.gguf

# jpt-4b: a merged LoRA, converted directly
scripts/build_jpt_gguf.sh <jpt-4b-dir> work/side/jpt-4b-q8_0.gguf

# neohorsejev-4b: convert the backbone, then export the pointer head
scripts/build_neohorsejev_gguf.sh <NeoHorse-Jev-4B-dir> work/side
python3 scripts/export_neohorsejev_head.py <NeoHorse-Jev-4B>/pointer_head.safetensors \
  work/side/neohorsejev-head.f32

Every export is quantized to Q8_0. kev merges the LoRA in fp32 before conversion and writes kev-head.f32 (the q and k pointer rows with their biases) plus kev.json. The same scripts handle the larger sizes with the matching base (Qwen3.5-2B for decider-2b, Qwen3.5-4B for kev-4b); pass --stream for a low-memory merge (one tensor at a time) on checkpoints too large to hold in RAM.

How it works

Dohnuts reads the post-norm hidden state at each candidate marker, applies one scalar head, and temperature-scales a softmax per question. llama.cpp supports the architecture (qwen35) and exposes all of it through the public API, so the core is unchanged:

Dohnuts Here
Prompt template with a reserved marker token per candidate protocol.cpp, tokenized without special tokens
Hidden state at each marker llama_get_embeddings_ith with embeddings=true, pooling off
Scalar scorer Linear(1024, 1) head.f32 dot product in engine.cpp
Temperature softmax, entropy confidence, expected score calibrate_answer in protocol.cpp
Shared input prefix common prefix decoded once, then llama_memory_seq_cp
Frozen vision tower + merger mtmd with mmproj-dohnuts-0.1.0-bf16.gguf; M-RoPE positions injected per image chunk

The side profiles use the same backend but their own readout: decider, thisthat and tev1 restrict the LM head to the option labels, and kev projects the decide and option-end hidden states through a bilinear pointer head. decider, thisthat and tev1 share the single-token label table and the letter-restricted softmax; thisthat also folds a request's questions into one prompt, one answer slot each, and reads them in one pass.

Norm weights are stored as weight + 1, matching the Dohnuts fused kernels.

Layout

include/dohnuts/engine.hpp    engine interface
include/dohnuts/protocol.hpp  prompt rendering, calibration, predictor
include/dohnuts/profile.hpp   side model profile switch
include/dohnuts/side.hpp      side engine facade
include/dohnuts/side/         runner, decider, thisthat, kev and tev1 profiles
include/dohnuts/http.hpp      HTTP transport
src/engine.cpp                llama.cpp/mtmd wrapper: load, tokenize, batched scoring
src/protocol.cpp              Dohnuts templates and answers
src/side.cpp                  side profile dispatch
src/side/                     shared runner and the four side profiles
src/http.cpp                  server routes and CORS
src/main.cpp                  CLI and HTTP server
scripts/                      setup, native/Windows builds, model export, GGUF
cmake/                        MinGW-w64 cross toolchain

License

The code is licensed under Apache-2.0. The HTTP layer is adapted from laya.cpp under MIT. See NOTICE for third-party terms. Dohnuts model weights keep their own license; the side models keep theirs.

About

The same decisions, on CPU. System One model that can run on your Personal Computer.

Topics

Resources

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages