Bunrin Bench, shown in our model reports, is a machine-scored benchmark that we built ourselves.
It measures only the things that tend to go wrong in real use: whether the model follows an output format, keeps numerical facts when summarizing, and does not turn hearsay into assertion.
All tasks use fictional material and the answers are in the prompt, so the test measures faithfulness to instructions rather than memorized knowledge.
Measurements are taken on our own machine (a MacBook Pro M5 Max with 128 GB of memory) using quantized builds (public versions made lighter at some cost in precision). Results vary with quantization and runtime, so each result names the build and environment tested.
The test may change over time. Scores from different test versions (v1 and so on) are not directly comparable.