I benchmarked seven local decision models on 22,270 questions
Seven open decision models disagree about what they are good at. I measured classification, moderation, scoring, latency, and memory, then published every probability needed to check the results.
#tools · #ai · #rust · #mlx · #benchmarks