==================================================================================================== PART 1. Per-passage probability: of passages given score s, what share were relevant? Queries: BM25 top-30 holds at least one relevant passage (same set eval.py uses). Unjudged passages count as not relevant. How thoroughly each dataset is labeled (queries used here): dataset queries labeled answers per query labeled answers in the 30 scifact 264 1.1 1.1 fiqa 411 2.8 1.6 nq 320 1.3 1.2 nfcorpus 239 48.2 5.5 trec-covid 50 493.5 17.3 bright-biology 39 4.0 1.6 bright-economics 38 12.2 3.1 csn-python 256 1.0 1.0 jev-noul-pair, all 8 English datasets pooled (n=48510, pooled ECE 0.082) bin n avg said share relevant 95% range 0.00-0.10 32965 0.032 0.013 0.012-0.014 0.10-0.20 4636 0.137 0.077 0.069-0.085 0.20-0.30 2218 0.241 0.120 0.107-0.135 0.30-0.40 1341 0.342 0.175 0.156-0.197 0.40-0.50 1243 0.445 0.245 0.222-0.270 0.50-0.60 1078 0.544 0.277 0.251-0.305 0.60-0.70 1106 0.645 0.306 0.279-0.333 0.70-0.80 1072 0.744 0.333 0.305-0.362 0.80-0.90 1189 0.847 0.392 0.365-0.420 0.90-1.00 1662 0.944 0.525 0.501-0.549 jev-noul-batch, all 8 English datasets pooled (n=48510, pooled ECE 0.075) bin n avg said share relevant 95% range 0.00-0.10 32754 0.039 0.010 0.009-0.011 0.10-0.20 5376 0.135 0.068 0.061-0.075 0.20-0.30 2372 0.241 0.148 0.135-0.163 0.30-0.40 1672 0.343 0.226 0.207-0.247 0.40-0.50 1258 0.443 0.282 0.258-0.308 0.50-0.60 1007 0.545 0.290 0.263-0.319 0.60-0.70 897 0.644 0.309 0.279-0.340 0.70-0.80 864 0.744 0.350 0.318-0.382 0.80-0.90 986 0.848 0.471 0.440-0.502 0.90-1.00 1324 0.941 0.617 0.591-0.643 deepseek-pair, all 8 English datasets pooled (n=48510, pooled ECE 0.097) bin n avg said share relevant 95% range 0.00-0.10 42250 0.000 0.028 0.026-0.030 0.90-1.00 6260 1.000 0.438 0.426-0.451 qwen-rlcd-pair, all 8 English datasets pooled (n=48510, pooled ECE 0.153) bin n avg said share relevant 95% range 0.00-0.10 15064 0.052 0.022 0.020-0.025 0.10-0.20 12076 0.146 0.050 0.046-0.054 0.20-0.30 7509 0.244 0.073 0.068-0.079 0.30-0.40 4960 0.347 0.107 0.099-0.116 0.40-0.50 2973 0.446 0.145 0.133-0.158 0.50-0.60 2190 0.543 0.187 0.171-0.204 0.60-0.70 1720 0.646 0.253 0.233-0.274 0.70-0.80 1058 0.749 0.266 0.240-0.293 0.80-0.90 701 0.846 0.354 0.319-0.390 0.90-1.00 259 0.933 0.382 0.325-0.443 qwen-rlcd-batch, all 8 English datasets pooled (n=48510, pooled ECE 0.238) bin n avg said share relevant 95% range 0.00-0.10 5043 0.066 0.037 0.032-0.043 0.10-0.20 10255 0.150 0.073 0.068-0.078 0.20-0.30 10121 0.247 0.090 0.085-0.096 0.30-0.40 8510 0.348 0.090 0.084-0.096 0.40-0.50 5572 0.445 0.085 0.078-0.092 0.50-0.60 4411 0.542 0.090 0.082-0.099 0.60-0.70 2557 0.646 0.093 0.083-0.105 0.70-0.80 1332 0.745 0.094 0.079-0.111 0.80-0.90 595 0.843 0.108 0.085-0.135 0.90-1.00 114 0.927 0.105 0.061-0.175 Cross-check, per-dataset ECE: this recount vs results/summary.json (should match to 3 decimals) jev-noul-pair mean of datasets 0.097 | scifact:0.023/0.023 fiqa:0.184/0.184 nq:0.139/0.139 nfcorpus:0.037/0.037 trec-covid:0.249/0.249 bright-biology:0.046/0.046 bright-economics:0.026/0.026 csn-python:0.072/0.072 jev-noul-batch mean of datasets 0.098 | scifact:0.032/0.032 fiqa:0.163/0.163 nq:0.129/0.129 nfcorpus:0.036/0.036 trec-covid:0.246/0.246 bright-biology:0.078/0.078 bright-economics:0.045/0.045 csn-python:0.058/0.058 deepseek-pair mean of datasets 0.112 | scifact:0.030/0.03 fiqa:0.124/0.124 nq:0.097/0.097 nfcorpus:0.170/0.17 trec-covid:0.271/0.271 bright-biology:0.071/0.071 bright-economics:0.115/0.115 csn-python:0.021/0.021 qwen-rlcd-pair mean of datasets 0.184 | scifact:0.139/0.139 fiqa:0.160/0.16 nq:0.205/0.205 nfcorpus:0.125/0.125 trec-covid:0.260/0.26 bright-biology:0.219/0.219 bright-economics:0.179/0.179 csn-python:0.183/0.183 qwen-rlcd-batch mean of datasets 0.240 | scifact:0.346/0.346 fiqa:0.133/0.133 nq:0.322/0.322 nfcorpus:0.147/0.147 trec-covid:0.237/0.237 bright-biology:0.218/0.218 bright-economics:0.131/0.131 csn-python:0.387/0.387 largest disagreement: 0.0000; raw-vs-stored score mismatches: {'jev-noul-pair': 0, 'jev-noul-batch': 0, 'deepseek-pair': 0, 'qwen-rlcd-pair': 0, 'qwen-rlcd-batch': 0} jev-noul-pair, trec-covid only (n=1500, pooled ECE 0.249) bin n avg said share relevant 95% range 0.00-0.10 545 0.041 0.202 0.170-0.238 0.10-0.20 238 0.138 0.634 0.572-0.693 0.20-0.30 124 0.239 0.702 0.616-0.775 0.30-0.40 85 0.338 0.835 0.742-0.899 0.40-0.50 60 0.449 0.700 0.575-0.801 0.50-0.60 61 0.539 0.770 0.651-0.858 0.60-0.70 70 0.650 0.886 0.790-0.941 0.70-0.80 70 0.745 0.857 0.757-0.921 0.80-0.90 110 0.846 0.927 0.863-0.963 0.90-1.00 137 0.942 0.985 0.948-0.996 jev-noul-pair, nq only (n=9600, pooled ECE 0.139) bin n avg said share relevant 95% range 0.00-0.10 6743 0.032 0.001 0.000-0.002 0.10-0.20 750 0.137 0.011 0.005-0.021 0.20-0.30 340 0.240 0.026 0.014-0.050 0.30-0.40 184 0.343 0.054 0.030-0.097 0.40-0.50 174 0.447 0.069 0.040-0.117 0.50-0.60 140 0.544 0.086 0.050-0.144 0.60-0.70 164 0.649 0.104 0.066-0.160 0.70-0.80 189 0.744 0.159 0.114-0.218 0.80-0.90 272 0.847 0.173 0.133-0.222 0.90-1.00 644 0.954 0.345 0.309-0.382 The 0.9-and-up bin only, per dataset, Jev yes/no per pair (worst first): jev-noul-pair: nq 34% (222/644), fiqa 46% (251/550), bright-biology 60% (9/15), nfcorpus 77% (59/77), scifact 79% (115/146), csn-python 87% (81/93), trec-covid 99% (135/137) jev-noul-batch: nq 42% (225/532), fiqa 58% (207/359), bright-biology 60% (12/20), bright-economics 67% (2/3), nfcorpus 82% (33/40), scifact 83% (121/146), csn-python 95% (122/128), trec-covid 99% (95/96) deepseek-pair: nq 26% (312/1185), fiqa 27% (541/1971), bright-biology 38% (31/82), bright-economics 42% (32/76), nfcorpus 53% (896/1699), csn-python 62% (249/403), scifact 64% (133/207), trec-covid 86% (549/637) ==================================================================================================== PART 2. Jev Choice: when it reports confidence c, how often is its pick right? ----- jev-choice ----- jev-choice: TypeSafe confidence field, lists WITH an answer (right = picked a relevant passage) (n=1617, pooled ECE 0.052) bin n avg said share right 95% range 0.00-0.50 345 0.378 0.409 0.358-0.461 0.50-0.60 171 0.544 0.532 0.457-0.605 0.60-0.70 156 0.644 0.692 0.616-0.759 0.70-0.80 152 0.746 0.651 0.573-0.723 0.80-0.90 189 0.851 0.704 0.635-0.764 0.90-0.95 138 0.921 0.855 0.787-0.904 0.95-0.99 214 0.968 0.921 0.876-0.950 0.99-1.00 252 0.995 0.984 0.960-0.994 jev-choice: TypeSafe confidence field, lists with NO answer (right = said none) (n=1609, pooled ECE 0.256) bin n avg said share right 95% range 0.00-0.50 488 0.371 0.287 0.249-0.329 0.50-0.60 197 0.544 0.330 0.268-0.398 0.60-0.70 177 0.645 0.373 0.305-0.446 0.70-0.80 170 0.746 0.359 0.291-0.433 0.80-0.90 218 0.849 0.500 0.434-0.566 0.90-0.95 126 0.923 0.532 0.445-0.617 0.95-0.99 157 0.966 0.631 0.553-0.702 0.99-1.00 76 0.993 0.500 0.390-0.610 jev-choice: TypeSafe confidence field, both kinds of list together (n=3226, pooled ECE 0.143) bin n avg said share right 95% range 0.00-0.50 833 0.374 0.337 0.306-0.370 0.50-0.60 368 0.544 0.424 0.374-0.475 0.60-0.70 333 0.645 0.523 0.469-0.576 0.70-0.80 322 0.746 0.497 0.443-0.551 0.80-0.90 407 0.850 0.595 0.546-0.641 0.90-0.95 264 0.922 0.701 0.643-0.753 0.95-0.99 371 0.967 0.798 0.754-0.836 0.99-1.00 328 0.995 0.872 0.831-0.904 jev-choice: probability Jev gave its own pick, both kinds together (n=3226, pooled ECE 0.161) bin n avg said share right 95% range 0.00-0.50 725 0.388 0.334 0.300-0.369 0.50-0.60 397 0.545 0.393 0.346-0.442 0.60-0.70 339 0.644 0.501 0.449-0.554 0.70-0.80 339 0.745 0.504 0.451-0.557 0.80-0.90 392 0.849 0.602 0.553-0.649 0.90-0.95 262 0.920 0.645 0.585-0.701 0.95-0.99 339 0.966 0.764 0.716-0.806 0.99-1.00 433 0.996 0.871 0.836-0.899 >> confidence 0.9 or more, answer in the list: right 563/604 = 93.2% (95% range 90.9% to 95.0%) >> confidence 0.9 or more, answer NOT in the list: right 204/359 = 56.8% (95% range 51.7% to 61.8%) >> confidence 0.9 or more, both kinds: right 767/963 = 79.6% (95% range 77.0% to 82.1%) >> per dataset, confidence 0.9 or more, worst first: bright-biology 41% (7/17), nq 60% (131/219), bright-economics 67% (4/6), trec-covid 75% (6/8), nfcorpus 77% (61/79), fiqa 80% (123/154), scifact 89% (232/260), csn-python 92% (203/220) cross-check vs summary.json thresholds (eval.py's rule: top-scored passage, lists with an answer): 0 band(s) differ eval.py's rule pooled over 8 datasets: 0.5-0.9 475/668 = 71.1%, <0.5 182/345 = 52.8%, >=0.9 572/604 = 94.7% ----- jev-choice-reversed ----- jev-choice-reversed: TypeSafe confidence field, lists WITH an answer (right = picked a relevant passage) (n=1617, pooled ECE 0.051) bin n avg said share right 95% range 0.00-0.50 378 0.372 0.418 0.369-0.468 0.50-0.60 168 0.545 0.571 0.496-0.644 0.60-0.70 160 0.644 0.606 0.529-0.679 0.70-0.80 142 0.746 0.683 0.603-0.754 0.80-0.90 183 0.849 0.721 0.652-0.781 0.90-0.95 137 0.923 0.832 0.761-0.885 0.95-0.99 199 0.968 0.940 0.898-0.965 0.99-1.00 250 0.995 0.980 0.954-0.991 jev-choice-reversed: TypeSafe confidence field, lists with NO answer (right = said none) (n=0, pooled ECE 0.000) bin n avg said share right 95% range jev-choice-reversed: TypeSafe confidence field, both kinds of list together (n=1617, pooled ECE 0.051) bin n avg said share right 95% range 0.00-0.50 378 0.372 0.418 0.369-0.468 0.50-0.60 168 0.545 0.571 0.496-0.644 0.60-0.70 160 0.644 0.606 0.529-0.679 0.70-0.80 142 0.746 0.683 0.603-0.754 0.80-0.90 183 0.849 0.721 0.652-0.781 0.90-0.95 137 0.923 0.832 0.761-0.885 0.95-0.99 199 0.968 0.940 0.898-0.965 0.99-1.00 250 0.995 0.980 0.954-0.991 jev-choice-reversed: probability Jev gave its own pick, both kinds together (n=1617, pooled ECE 0.050) bin n avg said share right 95% range 0.00-0.50 328 0.385 0.393 0.342-0.447 0.50-0.60 180 0.544 0.572 0.499-0.642 0.60-0.70 165 0.644 0.600 0.524-0.672 0.70-0.80 154 0.745 0.669 0.591-0.738 0.80-0.90 174 0.849 0.713 0.641-0.775 0.90-0.95 128 0.921 0.805 0.728-0.864 0.95-0.99 179 0.967 0.911 0.860-0.944 0.99-1.00 309 0.996 0.977 0.954-0.989 >> confidence 0.9 or more, answer in the list: right 546/586 = 93.2% (95% range 90.8% to 94.9%) >> confidence 0.9 or more, both kinds: right 546/586 = 93.2% (95% range 90.8% to 94.9%) >> per dataset, confidence 0.9 or more, worst first: bright-biology 80% (8/10), scifact 86% (119/138), nq 91% (99/109), nfcorpus 92% (33/36), fiqa 95% (93/98), csn-python 99% (190/191), bright-economics 100% (1/1), trec-covid 100% (3/3) cross-check vs summary.json thresholds (eval.py's rule: top-scored passage, lists with an answer): eval.py's rule pooled over 8 datasets: 0.5-0.9 466/653 = 71.4%, <0.5 187/378 = 49.5%, >=0.9 555/586 = 94.7% ----- jev-tournament ----- jev-tournament: TypeSafe confidence field, lists WITH an answer (right = picked a relevant passage) (n=1617, pooled ECE 0.077) bin n avg said share right 95% range 0.00-0.50 295 0.375 0.407 0.352-0.464 0.50-0.60 116 0.543 0.552 0.461-0.639 0.60-0.70 136 0.644 0.610 0.526-0.688 0.70-0.80 144 0.747 0.632 0.551-0.706 0.80-0.90 200 0.848 0.695 0.628-0.755 0.90-0.95 155 0.923 0.819 0.751-0.872 0.95-0.99 255 0.969 0.863 0.815-0.900 0.99-1.00 316 0.996 0.934 0.901-0.956 jev-tournament: TypeSafe confidence field, lists with NO answer (right = said none) (n=1609, pooled ECE 0.306) bin n avg said share right 95% range 0.00-0.50 442 0.364 0.294 0.254-0.338 0.50-0.60 182 0.543 0.302 0.240-0.372 0.60-0.70 177 0.646 0.305 0.242-0.376 0.70-0.80 172 0.746 0.326 0.260-0.399 0.80-0.90 205 0.845 0.371 0.308-0.439 0.90-0.95 127 0.921 0.504 0.418-0.589 0.95-0.99 172 0.967 0.605 0.530-0.675 0.99-1.00 132 0.994 0.447 0.365-0.532 jev-tournament: TypeSafe confidence field, both kinds of list together (n=3226, pooled ECE 0.185) bin n avg said share right 95% range 0.00-0.50 737 0.368 0.339 0.306-0.374 0.50-0.60 298 0.543 0.399 0.345-0.456 0.60-0.70 313 0.645 0.438 0.384-0.493 0.70-0.80 316 0.746 0.465 0.411-0.520 0.80-0.90 405 0.847 0.531 0.482-0.579 0.90-0.95 282 0.922 0.677 0.621-0.729 0.95-0.99 427 0.968 0.759 0.716-0.797 0.99-1.00 448 0.995 0.790 0.750-0.825 jev-tournament: probability Jev gave its own pick, both kinds together (n=3226, pooled ECE 0.227) bin n avg said share right 95% range 0.00-0.50 431 0.408 0.288 0.247-0.332 0.50-0.60 399 0.544 0.386 0.339-0.435 0.60-0.70 343 0.645 0.461 0.409-0.514 0.70-0.80 367 0.745 0.431 0.381-0.482 0.80-0.90 431 0.847 0.501 0.454-0.548 0.90-0.95 290 0.921 0.645 0.588-0.698 0.95-0.99 417 0.968 0.736 0.692-0.776 0.99-1.00 548 0.996 0.790 0.754-0.822 >> confidence 0.9 or more, answer in the list: right 642/726 = 88.4% (95% range 85.9% to 90.6%) >> confidence 0.9 or more, answer NOT in the list: right 227/431 = 52.7% (95% range 48.0% to 57.3%) >> confidence 0.9 or more, both kinds: right 869/1157 = 75.1% (95% range 72.5% to 77.5%) >> per dataset, confidence 0.9 or more, worst first: bright-biology 45% (10/22), bright-economics 50% (4/8), nq 54% (145/268), trec-covid 62% (10/16), nfcorpus 72% (86/119), fiqa 73% (146/199), scifact 87% (247/284), csn-python 92% (221/241) cross-check vs summary.json thresholds (eval.py's rule: top-scored passage, lists with an answer): 0 band(s) differ eval.py's rule pooled over 8 datasets: 0.5-0.9 413/596 = 69.3%, <0.5 154/295 = 52.2%, >=0.9 653/726 = 89.9% pick-to-passage mapping mismatches (should be 0): {'jev-choice': 0, 'jev-choice-reversed': 0, 'jev-tournament': 0}