I benchmarked seven local decision models on 22,270 questions
Seven open decision models disagree about what they are good at. I measured classification, moderation, scoring, latency, and memory, then published every probability needed to check the results.
- Published
- Reading time
- 13 min read
Two weeks ago I gave my terminal a 40-millisecond gut feeling. The first release of feelsneat ran three Laya checkpoints. It answered typed questions about text without generating text of its own.
The model list has since grown. feelsneat now runs four Decision 2.0 checkpoints, named Eos, Sol, Nox, and Lux. They range from 0.8 billion to 9 billion parameters and use a Qwen3.5 network with a decision head. All seven models execute locally through MLX on Apple Silicon.
That left me with a practical question. Which checkpoint should I use?
Parameter count did not answer it. Neither did one overall accuracy number. A support-routing classifier, a moderation probability, and a five-level quality score place different demands on a model. Latency and memory also vary by more than two orders of magnitude across the seven checkpoints.
I built an evaluation runner into feelsneat, assembled a public test suite, and ran every model through 6,118 source cases. The cases produced 22,270 decisions. I also measured a fixed warm performance profile with one, five, and ten questions per request.
The short result is that there is no best checkpoint. Decision 2.0 Eos is the strongest 77-way banking classifier in this test. Laya multilingual gives the best moderation probabilities. Decision 2.0 Lux gives the best ordered response-quality scores, but needs about 18 GB at the largest quick workload and takes 1.87 seconds for ten questions. The right model depends on the question.
The result table
All measurements below came from the same Apple M2 Ultra with a 76-core GPU and 128 GB of unified memory, running macOS 15.7.9. Laya used FP16. Decision 2.0 used its fixed mixed BF16 and FP32 contract.
Higher is better for BANKING77 macro-F1. Lower is better for Civil Comments Brier score and HelpSteer2 ranked probability score. Latency is warm in-process p50. Memory is the MLX peak during the ten-question workload.
| Model | BANKING77 macro-F1 | Civil Comments Brier | HelpSteer2 RPS | 1 question | 10 questions | Peak memory |
|---|---|---|---|---|---|---|
| Laya English | 0.361 | 0.0450 | 0.170 | 12.6 ms | 51.3 ms | 1,413 MiB |
| Laya multilingual | 0.334 | 0.0175 | 0.191 | 6.8 ms | 20.8 ms | 1,024 MiB |
| Laya typed-decisions | 0.364 | 0.0699 | 0.160 | 12.4 ms | 48.1 ms | 1,413 MiB |
| Decision 2.0 Eos, 0.8B | 0.838 | 0.0968 | 0.191 | 25.4 ms | 190.3 ms | 3,207 MiB |
| Decision 2.0 Sol, 2B | 0.800 | 0.1019 | 0.174 | 45.6 ms | 392.7 ms | 5,695 MiB |
| Decision 2.0 Nox, 4B | 0.812 | 0.1067 | 0.143 | 106.9 ms | 957.9 ms | 10,269 MiB |
| Decision 2.0 Lux, 9B | 0.833 | 0.1460 | 0.136 | 173.0 ms | 1,874.7 ms | 18,111 MiB |
This table is a map, not a ranking. Each quality column tests a different decision shape, dataset, prompt, and metric. The public data may also have appeared in model training. I call these numbers a public retest, not evidence of unseen generalization.
Eos handles 77 choices much better than Laya
BANKING77 contains short customer messages and 77 banking intents. The suite uses the complete official test split of 3,080 messages. Each case asks one question with all 77 intent labels as options.
That option count is deliberately uncomfortable for Laya. Its bidirectional encoder gives the question, every option, and the state one shared token budget. Upstream already warns that choice quality falls beyond about twenty options. The test makes that limit visible:
| Model | Accuracy | Macro-F1 |
|---|---|---|
| Laya English | 0.409 | 0.361 |
| Laya multilingual | 0.359 | 0.334 |
| Laya typed-decisions | 0.408 | 0.364 |
| Decision 2.0 Eos | 0.842 | 0.838 |
| Decision 2.0 Sol | 0.805 | 0.800 |
| Decision 2.0 Nox | 0.817 | 0.812 |
| Decision 2.0 Lux | 0.837 | 0.833 |
Eos and Lux are only 0.0048 apart on macro-F1. The smallest Decision 2.0 model has the higher point estimate. Sol and Nox sit below both of them. More parameters do not produce a monotonic improvement here.
The result also tells me when not to reach for Laya. Its speed is attractive, but a flat classifier with 77 options is outside its comfortable range. I would choose Eos for this shape before spending another 15 GB on Lux.
The small multilingual encoder wins moderation
The Civil Comments portion uses the first 2,000 rows of the official test split. Selection does not inspect the labels. Each comment gets seven yes-or-no questions covering toxicity, severe toxicity, obscenity, threats, insults, identity attacks, and sexually explicit content. The targets preserve the source crowd fractions instead of turning every label into a hard zero or one.
Brier score is useful here because it measures the probability, not only which side of 0.5 won. The pooled scores over 14,000 decisions were:
| Model | Brier | AUROC |
|---|---|---|
| Laya English | 0.0450 | 0.9815 |
| Laya multilingual | 0.0175 | 0.9945 |
| Laya typed-decisions | 0.0699 | 0.9768 |
| Decision 2.0 Eos | 0.0968 | 0.8594 |
| Decision 2.0 Sol | 0.1019 | 0.8668 |
| Decision 2.0 Nox | 0.1067 | 0.9267 |
| Decision 2.0 Lux | 0.1460 | 0.9458 |
Laya multilingual is both the smallest checkpoint in the catalog and the clear winner on this test. Lux has the worst Brier score despite having the most parameters.
I would still not ship a moderation threshold from this table. Civil Comments is public, its labels represent subjective judgments, and the seven attributes have different class balances. A production threshold needs local examples, the actual cost of false positives, and a held-out set that was not used while choosing the prompt. The result only says which model I would test first.
Lux helps on ordered scores
HelpSteer2 supplies a prompt, a response, and five labels from zero to four: helpfulness, correctness, coherence, complexity, and verbosity. The suite uses all 1,038 rows from its validation split, which produces 5,190 score decisions.
For an ordered label, predicting level three when the answer is four is less wrong than predicting zero. Ranked probability score measures that distance across the complete distribution. feelsneat normalizes it by the number of level boundaries, so zero is perfect and one is worst.
| Model | Mean absolute error | Ranked probability score |
|---|---|---|
| Laya English | 0.996 | 0.170 |
| Laya multilingual | 1.056 | 0.191 |
| Laya typed-decisions | 0.983 | 0.160 |
| Decision 2.0 Eos | 1.074 | 0.191 |
| Decision 2.0 Sol | 1.019 | 0.174 |
| Decision 2.0 Nox | 0.879 | 0.143 |
| Decision 2.0 Lux | 0.845 | 0.136 |
This is the one task where the two largest Decision 2.0 checkpoints justify part of their cost. Nox closes most of the gap while using about 8 GB less peak memory than Lux. Laya typed-decisions remains a reasonable low-latency option, especially when a one-level error is acceptable.
Complexity and verbosity deserve special care. They describe the response rather than whether the response is good. Treating every high score as a positive outcome would change the meaning of the source labels.
Why I did not publish one overall winner
The combined report contains an overall accuracy field because every decision has a predicted label. That number is easy to misuse.
Civil Comments contributes 14,000 of the 22,270 decisions and most labels are negative. A model can raise overall accuracy by becoming a better hard-threshold moderation classifier while becoming worse at banking intents and response scores. The seven moderation questions also come from the same comment, so they are not seven independent source examples.
For those reasons, feelsneat reports each task and question separately. Its paired comparison groups decisions by source case before bootstrapping. Failures stay in the accuracy denominator. Probability metrics use only scored rows and always report how many rows failed.
I prefer the table of trade-offs. It preserves the question I actually need to answer: how much quality, latency, and memory does a checkpoint provide for my decision shape?
Performance changes the recommendation
The quick performance profile uses one support-ticket state and a fixed mix of yes-or-no, choice, and score questions. Each workload gets three warmups and twenty timed samples. The report includes rendering, tokenization, MLX execution, and synchronization. It excludes model download, cold process startup, HTTP queueing, and energy use.
The Laya models benefit sharply from batching. English takes 12.6 ms for one question and 51.3 ms for ten, or 5.1 ms per question. Multilingual reaches 481 questions per second in the ten-question workload.
Decision 2.0 processes a separate causal sequence for each question. Its cost grows much closer to linearly. Eos remains comfortable at 25.4 ms for one question and 190.3 ms for ten. Nox reaches 957.9 ms for ten, while Lux takes 1.87 seconds. Lux’s ten-question p95 was 1.97 seconds.
Memory places a harder boundary. Eos peaked at 3.2 GB in the largest quick workload. Sol used 5.7 GB, Nox 10.3 GB, and Lux 18.1 GB. Nox leaves little room for another large local model on a 16 GB Mac. Lux needs a larger-memory machine.
The complete quality run took 7 hours and 34 seconds wall time. The model inference totals were 2 minutes 53 seconds for multilingual, 9 minutes 19 seconds for typed-decisions, 11 minutes 36 seconds for English, 25 minutes 58 seconds for Eos, 46 minutes 50 seconds for Sol, 1 hour 54 minutes 47 seconds for Nox, and 3 hours 27 minutes 49 seconds for Lux. BANKING77 is expensive for Decision 2.0 because every request contains 77 rendered candidates.
These timings came from one machine on one run. Small differences can move with thermals, operating-system activity, and MLX versions. The raw performance report records the workload digest, model artifact hashes, actual token counts, p50, p95, p99, memory counters, and environment so another run can be compared honestly.
The report keeps every probability
A benchmark summary is not enough to audit a probability model. I wanted to answer questions after the run without loading 38 GB of checkpoints again.
The evaluation report therefore stores every case ID, expected answer, predicted answer, unrounded probability vector, error, confidence, token count, request time, model revision, and artifact checksum. Every aggregate in this article can be recalculated from that file.
That completeness has a cost. The raw JSON is 185 MB because BANKING77 alone stores 77 probabilities for every model and case. It compresses to 29 MB. I published the full compressed report, a smaller aggregate summary, the performance report, and their SHA-256 checksums.
The evaluation bundle has digest b4cd29250bb20a59a73eb52a923cf10033445e95c2b6be5aa2db737402120bcc. Changing a prompt, option, case, or source revision changes the digest. feelsneat refuses a paired comparison between different digests.
Replay the small version first
feelsneat currently needs an Apple Silicon Mac and Rust. Installation builds MLX, so the first compile takes several minutes:
cargo install --locked --git https://gitlab.com/parlant-co/feelsneat.git
Eos is the sensible first Decision 2.0 checkpoint. It needs a 2.04 GB download and performs well on the wide choice task:
feelsneat --setup --model decision2-eos
feelsneat eval setup feelsneat-core-public-v1
feelsneat eval validate feelsneat-core-public-v1
feelsneat --model decision2-eos eval run feelsneat-core-public-v1 \
--task banking77 --limit 100 --output eos-smoke.json
Remove --limit 100 for the complete BANKING77 split. The task filter and limit become part of a derived digest, so a smoke report cannot masquerade as the full suite.
The quick performance profile takes less than a minute for Eos:
feelsneat --model decision2-eos bench \
--profile quick --runs 20 --output eos-performance.json
Use the standard profile before publishing close latency comparisons. It runs more samples and adds state lengths up to about 8,192 words, choices with 4, 20, and 77 options, and batches of 25 questions:
feelsneat --model decision2-eos bench \
--profile standard --output eos-standard.json
Replay all seven models
The complete run needs all checkpoints, about 38 GB of downloads, enough free disk space for verified copies, and enough unified memory for Lux:
feelsneat --setup --model all
feelsneat eval setup feelsneat-core-public-v1
feelsneat --model all eval run feelsneat-core-public-v1 \
--output core-public-all.json
feelsneat --model all bench --profile quick --runs 20 \
--output performance-quick-all.json
Models run sequentially and unload between runs. On my M2 Ultra, the evaluation command occupied the machine for seven hours. Start it in tmux if you intend to reproduce the full file.
To inspect the published artifacts without running a model:
base=https://gitlab.com/api/v4/projects/86892675/packages/generic/benchmarks/2026-10-08
for file in core-public-all.json.gz core-public-summary.json \
performance-quick-all.json SHA256SUMS; do
curl -fLO "$base/$file"
done
shasum -a 256 -c SHA256SUMS
gzip -dk core-public-all.json.gz
jq '.runs[] | {model, metrics: .metrics.overall}' core-public-all.json
The evaluation documentation defines every field and metric. The deterministic suite builder pins the three source revisions and never executes downloaded Python packages or dataset repository code.
Test the decisions you actually make
A public benchmark helped me find the broad limits. It cannot tell me whether a model understands my support queues, my safety policy, or the way my users write.
feelsneat can create a local evaluation with the same request shape as its normal API:
feelsneat eval init my-eval
$EDITOR my-eval/cases.jsonl
feelsneat eval validate my-eval/cases.jsonl
feelsneat --model english,decision2-eos,decision2-nox \
eval run my-eval --output my-results.json
Each JSONL case contains a state, typed questions, expected answers, and optional tags. Keep private cases outside Git. Reserve a holdout before changing prompts. For subjective questions, preserve the individual votes instead of erasing disagreement with one hard label.
I want more evidence from real workflows. Install feelsneat, try Eos for wide choices, try Laya multilingual for fast binary judgments, and measure the task that matters to you. If a model fails in an interesting way, open an issue with the request shape, model slug, and aggregate result. Do not attach private source text.
The useful outcome is not a larger leaderboard. It is knowing which local model deserves a narrow job, how often code should trust it, and when the decision must go to a person or a stronger reasoning model.