Bench results¶
Labels are the neon-avu-eval silver (agreement between four agentic models, not truth;
see Learning). L1 (LOCO) cells are pooled held-out decisions
over seven leave-one-card-out folds. cov@5% is the share of items that can be answered
before the error rate on the answered set exceeds 5%; flip is the answer change under
option reversal (choice) or phrasing swap (yes/no). Every number below names its JSON.
End to end: the whole graph, twice per card (carc-fast at L0), 2026-09-25¶
bench e2e runs the decision graph on the seven fixture cards, two repetitions per planner,
and reports agreement rather than asserting determinism: the Jaccard overlap of proposed AVU
triples between repetitions, and the recall of the neon-avu-eval consensus terms (proposed by
every agentic model, or by at least two). Source: bench/results/2026-09-25/RedHatAI__Qwen3-8B-NVFP4.gateway.e2e.json.
| planner | rep agreement | consensus-all recall | consensus-majority recall | fallbacks |
|---|---|---|---|---|
| static | 0.37 | 0.25 | 0.06 | 0 / 7 |
| gateway (carc-tools) | 0.18 | 0.05 | 0.05 | 1 / 7 |
At L0 on carc-fast the graph is far from repeatable (per card 0.0 to 0.67) and recovers
little of the agentic consensus; proposals per card were 1 to 9. The gateway planner made it
worse and cost 96 minutes of wall time at ~4.7 tokens per second. The planner audit
(dataset.ontology_applies) rejected gaz, pato, uo and ro on most cards where the
static planner had named them. Specificity opened 23 child groups and replaced no parent.
These are the numbers the L2 head and curator labels have to move.
Qwen/Qwen3-8B bf16 on the GB10 workstation (local weights), 2026-09-25¶
The same labels through AnyJev's HFBackend on this host: the full vocabulary is read (no
top-20 cap, so no erased labels), and L2 heads read the hidden state of one block.
L1 (LOCO) and L2 (LOCO) cells are pooled held-out decisions over seven leave-one-card-out
folds. The prompts column reads 0 because the local backend had no request counter in this
run (added afterwards; the next run reports prefills).
| task | n | level | acc | ece | brier | cov@5% | cov@10% | flip | n_neg | prompts | ms/dec | | neon_annotate | 98 | L0 | 0.439 | 0.474 | 0.970 | 0.000 | 0.000 | 0.143 | 35 | 0 | 817.1 | | neon_annotate | 98 | L1 (LOCO) | 0.294 | 0.529 | 0.569 | 0.000 | 0.000 | | 5 | 0 | 369.2 | | neon_annotate | 98 | L2 (LOCO) | 0.765 | 0.257 | 0.354 | 0.353 | 0.353 | | 5 | 0 | 196.5 | | neon_annotate | 98 | raw | 0.439 | 0.512 | 1.022 | 0.000 | 0.000 | | 35 | 0 | 398.8 | | neon_aspect | 60 | L0 | 0.533 | 0.456 | 0.922 | 0.000 | 0.000 | 0.150 | 0 | 0 | 1917.5 | | neon_aspect | 60 | L1 (LOCO) | 0.517 | 0.162 | 0.633 | 0.000 | 0.000 | | 0 | 0 | 8175.3 | | neon_aspect | 60 | L2 (LOCO) | 0.567 | 0.175 | 0.552 | 0.383 | 0.417 | | 0 | 0 | 2983.5 | | neon_aspect | 60 | raw | 0.550 | 0.420 | 0.841 | 0.117 | 0.117 | 0.250 | 0 | 0 | 447.3 | | neon_keep_avu | 313 | L0 | 0.949 | 0.031 | 0.085 | 0.994 | 1.000 | 0.022 | 11 | 0 | 845.2 | | neon_keep_avu | 313 | L1 | n/a | | | | | | | | | | neon_keep_avu | 313 | L2 | n/a | | | | | | | | | | neon_keep_avu | 313 | raw | 0.962 | 0.036 | 0.074 | 1.000 | 1.000 | | 11 | 0 | 744.1 | | neon_ontology_fits | 190 | L0 | 0.426 | 0.560 | 1.111 | 0.089 | 0.116 | 0.084 | 114 | 0 | 621.4 | | neon_ontology_fits | 190 | L1 (LOCO) | 0.426 | 0.233 | 0.537 | 0.089 | 0.111 | | 114 | 0 | 5381.1 | | neon_ontology_fits | 190 | L2 (LOCO) | 0.837 | 0.078 | 0.231 | 0.547 | 0.768 | | 114 | 0 | 1906.4 | | neon_ontology_fits | 190 | raw | 0.421 | 0.566 | 1.137 | 0.089 | 0.105 | | 114 | 0 | 216.5 | | neon_ontology_for_column | 59 | L0 | 0.441 | 0.508 | 1.027 | 0.169 | 0.271 | 0.186 | 0 | 0 | 3085.4 | | neon_ontology_for_column | 59 | L1 (LOCO) | 0.441 | 0.141 | 0.732 | 0.102 | 0.102 | | 0 | 0 | 6622.9 | | neon_ontology_for_column | 59 | L2 (LOCO) | 0.712 | 0.154 | 0.391 | 0.458 | 0.678 | | 0 | 0 | 1679.8 | | neon_ontology_for_column | 59 | raw | 0.492 | 0.443 | 0.896 | 0.051 | 0.051 | 0.136 | 0 | 0 | 568.5 | | neon_term_choice26 | 44 | L0 | 0.114 | 0.833 | 1.648 | 0.000 | 0.000 | 0.236 | 28 | 0 | 578.2 | | neon_term_choice26 | 44 | L1 | n/a | | | | | | | | | | neon_term_choice26 | 44 | L2 | n/a | | | | | | | | | | neon_term_choice26 | 44 | raw | 0.114 | 0.789 | 1.621 | 0.000 | 0.000 | | 28 | 0 | 191.4 | | neon_term_fits | 285 | L0 | 0.663 | 0.293 | 0.616 | 0.077 | 0.182 | 0.119 | 199 | 0 | 392.3 | | neon_term_fits | 285 | L1 (LOCO) | 0.656 | 0.079 | 0.422 | 0.049 | 0.077 | | 199 | 0 | 2765.1 | | neon_term_fits | 285 | L2 (LOCO) | 0.765 | 0.058 | 0.327 | 0.088 | 0.396 | | 199 | 0 | 1374.7 | | neon_term_fits | 285 | raw | 0.642 | 0.316 | 0.656 | 0.074 | 0.112 | | 199 | 0 | 199.1 | | neon_value_kind | 278 | L0 | 0.421 | 0.477 | 1.022 | 0.014 | 0.014 | 0.385 | 0 | 0 | 1046.0 | | neon_value_kind | 278 | L1 (LOCO) | 0.428 | 0.088 | 0.669 | 0.007 | 0.007 | | 0 | 0 | 1709.4 | | neon_value_kind | 278 | L2 (LOCO) | 0.576 | 0.069 | 0.565 | 0.029 | 0.040 | | 0 | 0 | 1275.0 | | neon_value_kind | 278 | raw | 0.406 | 0.536 | 1.127 | 0.014 | 0.014 | 0.500 | 0 | 0 | 370.4 |
How to read it (bench/results/2026-09-25/Qwen__Qwen3-8B.hf.json):
- term.fits (285 states, 199 negatives): accuracy 0.642 raw, 0.663 L0, 0.656 L1, 0.765
L2; ECE 0.316, 0.293, 0.079, 0.058; coverage at 5% risk 0.074, 0.077, 0.049, 0.088 and at
10% risk 0.112, 0.182, 0.077, 0.396. L2 meets the 0.10 ECE bar with the best accuracy,
but only 25 of 285 items can be answered at 5% risk, so
term.fitsautostays null. - column.ontology_fits (190, 114 negatives) is the one question where L2 changes the
picture: accuracy 0.837 and ECE 0.078 with coverage 0.547 at 5% risk and 0.768 at 10%,
against 0.426 / 0.233 / 0.089 at L1. The
policy_suggestionsblock cites this cell. Before any threshold is adopted the number needs a second label source: these labels are the agentic models' ontology choices, so the head has learnt their habits, not the truth. - column.annotate (98, 5 negatives above the weight cut at L1/L2): L2 accuracy 0.765 against 0.294 at L1, but with five negatives the coverage figure (0.353) rests on almost nothing and every fold barely passes the class guard.
- column.aspect (60) and avu.value_kind (278) gain accuracy at L2 (0.567, 0.576) and calibrate (ECE 0.175, 0.069) but stay near zero coverage at 5% risk.
- neon_term_choice26 at raw/L0 scores 0.114 locally against 0.227 / 0.205 on the gateway: reading the full vocabulary does not rescue a 26-way choice with a new key per item; the noul-per-candidate design is the right one for this model.
- avu.keep: 11 negatives, no fold may fit, unchanged.
- Against the gateway at L0 the local bf16 checkpoint scores term.fits 0.663 vs 0.611 and annotate 0.439 vs 0.398, with flip rates 0.119 vs 0.393 on term.fits: the NVFP4 quant and the bf16 weights are not interchangeable, which is why bundles are keyed per model (D5).
carc-fast (RedHatAI/Qwen3-8B-NVFP4) through the CARC gateway, 2026-09-25¶
| task | n | level | acc | ece | brier | cov@5% | cov@10% | flip | n_neg | prompts | ms/dec |
|---|---|---|---|---|---|---|---|---|---|---|---|
| neon_term_fits | 285 | raw | 0.502 | 0.372 | 0.780 | 0.004 | 0.004 | 199 | 285 | 100.6 | |
| neon_term_fits | 285 | L0 | 0.611 | 0.262 | 0.605 | 0.007 | 0.049 | 0.393 | 199 | 570 | 126.5 |
| neon_term_fits | 285 | L1 (LOCO) | 0.621 | 0.072 | 0.459 | 0.007 | 0.042 | 199 | 3990 | 728.9 | |
| neon_term_choice26 | 44 | raw | 0.227 | 0.689 | 1.443 | 0.000 | 0.000 | 28 | 44 | 151.6 | |
| neon_term_choice26 | 44 | L0 | 0.205 | 0.689 | 1.460 | 0.000 | 0.000 | 0.343 | 28 | 152 | 354.5 |
| neon_term_choice26 | 44 | L1 | n/a | ||||||||
| neon_annotate | 98 | raw | 0.378 | 0.517 | 1.062 | 0.000 | 0.000 | 35 | 98 | 46.8 | |
| neon_annotate | 98 | L0 | 0.398 | 0.490 | 1.020 | 0.000 | 0.000 | 0.316 | 35 | 196 | 39.1 |
| neon_annotate | 98 | L1 (LOCO) | 0.294 | 0.467 | 0.560 | 0.000 | 0.000 | 5 | 196 | 40.4 | |
| neon_aspect | 60 | raw | 0.600 | 0.376 | 0.775 | 0.000 | 0.000 | 0.433 | 0 | 120 | 83.2 |
| neon_aspect | 60 | L0 | 0.583 | 0.373 | 0.736 | 0.000 | 0.000 | 0.283 | 0 | 444 | 242.0 |
| neon_aspect | 60 | L1 (LOCO) | 0.600 | 0.181 | 0.552 | 0.017 | 0.017 | 0 | 1680 | 718.0 | |
| neon_ontology_fits | 190 | raw | 0.411 | 0.575 | 1.140 | 0.047 | 0.063 | 114 | 190 | 47.4 | |
| neon_ontology_fits | 190 | L0 | 0.474 | 0.417 | 0.868 | 0.032 | 0.116 | 0.289 | 114 | 380 | 47.6 |
| neon_ontology_fits | 190 | L1 (LOCO) | 0.458 | 0.182 | 0.494 | 0.068 | 0.105 | 114 | 2660 | 310.7 | |
| neon_keep_avu | 313 | raw | 0.965 | 0.025 | 0.064 | 1.000 | 1.000 | 11 | 313 | 171.9 | |
| neon_keep_avu | 313 | L0 | 0.895 | 0.044 | 0.160 | 0.786 | 0.974 | 0.323 | 11 | 626 | 178.4 |
| neon_keep_avu | 313 | L1 | n/a | ||||||||
| neon_value_kind | 278 | raw | 0.507 | 0.457 | 0.937 | 0.004 | 0.004 | 0.259 | 0 | 556 | 73.0 |
| neon_value_kind | 278 | L0 | 0.475 | 0.462 | 0.942 | 0.004 | 0.004 | 0.194 | 0 | 1578 | 124.2 |
| neon_value_kind | 278 | L1 (LOCO) | 0.468 | 0.124 | 0.642 | 0.004 | 0.004 | 0 | 4759 | 377.2 |
How to read it (bench/results/2026-09-25/RedHatAI__Qwen3-8B-NVFP4.gateway.json):
- term.fits (285 candidate states, 199 negatives): accuracy 0.502 raw, 0.611 at L0,
0.621 at L1; ECE 0.372, 0.262, 0.072. Calibration meets the pre-registered 0.10 bar at
L1, but coverage at 5% risk is 0.007 (two items), so every
autothreshold stays null. - neon_term_choice26 is the control: one choice per column over its candidates with a new key per item, so it can only run raw/L0 and never learns (accuracy 0.23 / 0.21).
- column.aspect (K=8) already lost 19 labels to the gateway's top-20 readout at L0; the wider the choice, the more the gateway erases (DESIGN D11).
- avu.keep has only 11 negatives, so no leave-one-card-out fold is allowed to fit; its raw/L0 numbers reflect a positives-heavy label set, not a calibrated gate.
- In M2 the bench and
learn fitdisagreed on term.fits ECE (0.072 here, 0.231 in the bundle's manifest). The cause was not the fitting:learn fitaveraged per-fold ECE and coverage, which are not linear in the items, while the bench pooled the held-out decisions. Since M3 both share one fitting path and both pool (DESIGN D18).
Synthetic backend (planted biases; numbers are not a claim about any model)¶
| task | n | level | acc | ece | brier | cov@5% | cov@10% | flip | n_neg | prompts | ms/dec |
|---|---|---|---|---|---|---|---|---|---|---|---|
| neon_term_fits | 285 | raw | 0.442 | 0.504 | 1.017 | 0.007 | 0.007 | 199 | 285 | 0.1 | |
| neon_term_fits | 285 | L0 | 0.442 | 0.488 | 0.946 | 0.007 | 0.007 | 0.004 | 199 | 570 | 0.2 |
| neon_term_fits | 285 | L1 (LOCO) | 0.442 | 0.178 | 0.508 | 0.007 | 0.007 | 199 | 3990 | 1.2 | |
| neon_term_choice26 | 44 | raw | 0.250 | 0.249 | 0.811 | 0.000 | 0.000 | 28 | 44 | 0.2 | |
| neon_term_choice26 | 44 | L0 | 0.159 | 0.317 | 0.841 | 0.000 | 0.000 | 0.411 | 28 | 221 | 0.7 |
| neon_term_choice26 | 44 | L1 | n/a | ||||||||
| neon_annotate | 98 | raw | 0.643 | 0.238 | 0.573 | 0.020 | 0.020 | 35 | 98 | 0.1 | |
| neon_annotate | 98 | L0 | 0.622 | 0.101 | 0.473 | 0.020 | 0.020 | 0.163 | 35 | 196 | 0.2 |
| neon_annotate | 98 | L1 (LOCO) | 0.647 | 0.353 | 0.457 | 0.118 | 0.118 | 5 | 196 | 0.1 | |
| neon_aspect | 60 | raw | 0.200 | 0.216 | 0.964 | 0.017 | 0.017 | 0.917 | 0 | 120 | 0.3 |
| neon_aspect | 60 | L0 | 0.117 | 0.156 | 0.878 | 0.000 | 0.000 | 0.000 | 0 | 861 | 1.8 |
| neon_aspect | 60 | L1 (LOCO) | 0.167 | 0.164 | 0.890 | 0.000 | 0.000 | 0 | 3074 | 6.3 | |
| neon_ontology_fits | 190 | raw | 0.400 | 0.480 | 0.938 | 0.000 | 0.000 | 114 | 190 | 0.1 | |
| neon_ontology_fits | 190 | L0 | 0.432 | 0.167 | 0.549 | 0.000 | 0.000 | 0.189 | 114 | 380 | 0.2 |
| neon_ontology_fits | 190 | L1 (LOCO) | 0.432 | 0.123 | 0.501 | 0.000 | 0.000 | 114 | 2660 | 1.0 | |
| neon_keep_avu | 313 | raw | 0.965 | 0.085 | 0.083 | 1.000 | 1.000 | 11 | 313 | 0.2 | |
| neon_keep_avu | 313 | L0 | 0.911 | 0.318 | 0.352 | 0.888 | 1.000 | 0.173 | 11 | 626 | 0.3 |
| neon_keep_avu | 313 | L1 | n/a | ||||||||
| neon_value_kind | 278 | raw | 0.478 | 0.130 | 0.691 | 0.007 | 0.007 | 1.000 | 0 | 556 | 0.3 |
| neon_value_kind | 278 | L0 | 0.216 | 0.112 | 0.757 | 0.004 | 0.004 | 0.007 | 0 | 2195 | 0.8 |
| neon_value_kind | 278 | L1 (LOCO) | 0.227 | 0.107 | 0.750 | 0.004 | 0.004 | 0 | 7661 | 2.4 |