============================================================================================================== PART A. One relevance question per passage, eleven wordings in one request. 1617 questions, 48510 passages. Model: ['jev-1.13.0'] share of passages labeled relevant: 0.081 wording avg said ECE all ECE trec >=0.9 share >=0.9 relevant AUROC top1 nDCG10 noul 0.155 0.075 0.247 2.8% 61.4% (59% to 64%) 0.906 74.0% 0.736 noul-dup 0.156 0.075 0.246 2.8% 61.4% (59% to 64%) 0.906 74.5% 0.738 yesno 0.118 0.059 0.274 5.3% 53.5% (52% to 55%) 0.888 73.3% 0.729 yesno-dup 0.118 0.059 0.276 5.3% 53.5% (52% to 55%) 0.887 73.5% 0.729 yesno-rev 0.123 0.062 0.255 5.2% 53.4% (51% to 55%) 0.890 73.7% 0.730 ab 0.117 0.056 0.276 4.9% 55.2% (53% to 57%) 0.888 73.1% 0.729 ab-rev 0.127 0.061 0.249 5.2% 54.1% (52% to 56%) 0.892 73.8% 0.733 xy 0.116 0.042 0.292 3.4% 59.6% (57% to 62%) 0.880 74.2% 0.733 xy-rev 0.176 0.095 0.236 4.3% 56.8% (55% to 59%) 0.864 71.7% 0.713 relevant 0.147 0.076 0.211 6.7% 50.5% (49% to 52%) 0.901 73.6% 0.734 relevant-rev 0.158 0.085 0.188 7.2% 49.4% (48% to 51%) 0.901 73.7% 0.733 choice-avg 0.135 0.055 0.245 4.9% 55.2% (53% to 57%) 0.888 72.9% 0.725 per-dataset ECE (worst first is not meaningful here; same order as the benchmark): noul scifact 0.032 fiqa 0.163 nq 0.129 nfcorpus 0.037 covid 0.247 biology 0.078 economics 0.044 python 0.058 yesno scifact 0.019 fiqa 0.123 nq 0.100 nfcorpus 0.090 covid 0.274 biology 0.061 economics 0.060 python 0.018 ab scifact 0.018 fiqa 0.122 nq 0.103 nfcorpus 0.095 covid 0.276 biology 0.057 economics 0.062 python 0.019 xy scifact 0.014 fiqa 0.116 nq 0.104 nfcorpus 0.091 covid 0.292 biology 0.048 economics 0.056 python 0.030 relevant scifact 0.022 fiqa 0.165 nq 0.120 nfcorpus 0.065 covid 0.211 biology 0.074 economics 0.059 python 0.027 How far two wordings sit apart, per passage (first minus second) pair mean diff mean |diff| |diff|>=0.10 |diff|>=0.30 noul vs noul-dup (exact duplicate (noise)) -0.000 0.005 0.0% 0.0% yesno vs yesno-dup (exact duplicate (noise)) -0.000 0.005 0.1% 0.0% yesno vs yesno-rev (criteria order) -0.005 0.009 1.5% 0.0% ab vs ab-rev (criteria order) -0.010 0.012 2.5% 0.0% xy vs xy-rev (criteria order) -0.060 0.061 17.5% 0.0% relevant vs relevant-rev (criteria order) -0.011 0.013 3.2% 0.0% ab vs yesno (labels) -0.001 0.009 1.0% 0.0% xy vs yesno (labels) -0.002 0.027 7.0% 0.1% relevant vs yesno (labels) +0.029 0.029 11.3% 0.6% xy vs ab (labels) -0.001 0.021 4.1% 0.0% yesno vs noul (Choice vs Noul) -0.038 0.055 14.8% 0.0% ab vs noul (Choice vs Noul) -0.038 0.053 12.2% 0.0% xy vs noul (Choice vs Noul) -0.039 0.045 9.5% 0.0% relevant vs noul (Choice vs Noul) -0.009 0.057 11.9% 0.4% Noul here vs the published 30-yes/no run (a separate request a day earlier): mean |diff| 0.006 over 1617 questions published 30-yes/no vs the same request sent again today: mean |diff| 0.006 Reproducibility: the same eleven-wording request sent again (200 questions) wording minus yesno, 1st minus yesno, 2nd same wording 1st vs 2nd, mean |diff| noul +0.041 +0.041 0.006 noul-dup +0.041 +0.041 0.006 yesno +0.000 +0.000 0.006 yesno-dup +0.000 +0.000 0.006 yesno-rev +0.007 +0.007 0.007 ab -0.002 -0.002 0.006 ab-rev +0.010 +0.010 0.007 xy -0.004 -0.004 0.007 xy-rev +0.059 +0.059 0.014 relevant +0.032 +0.032 0.007 relevant-rev +0.046 +0.045 0.008 per passage, does the xy-rev minus xy shift repeat? correlation between the two requests: 0.87 Choice minus Noul by dataset (does the direction change?): scifact yesno -0.030 ab -0.029 xy -0.027 relevant -0.022 fiqa yesno -0.040 ab -0.041 xy -0.047 relevant +0.002 nq yesno -0.029 ab -0.026 xy -0.025 relevant -0.009 nfcorpus yesno -0.049 ab -0.060 xy -0.066 relevant +0.002 trec-covid yesno -0.028 ab -0.029 xy -0.046 relevant +0.043 bright-biology yesno -0.042 ab -0.043 xy -0.042 relevant -0.014 bright-economics yesno -0.058 ab -0.060 xy -0.063 relevant -0.020 csn-python yesno -0.039 ab -0.038 xy -0.028 relevant -0.030 Where the wording matters: passages grouped by the published 30-yes/no probability (a separate earlier request, so the grouping is not tied to these answers) published P passages relevant label spread order gap duplicate gap spread>=0.3 0.0 to 0.1 32754 1.0% 0.009 0.014 0.000 0.0% 0.1 to 0.3 7748 9.2% 0.074 0.038 0.008 0.5% 0.3 to 0.7 4834 26.9% 0.254 0.070 0.025 32.9% 0.7 to 0.9 1850 41.4% 0.160 0.030 0.010 6.0% 0.9 to 1.0 1324 61.7% 0.032 0.005 0.001 0.0% Do the wordings agree on the top passage? (top by probability, ties keep BM25 order) noul-dup same top passage as noul on 95.7% yesno same top passage as noul on 89.7% yesno-dup same top passage as noul on 89.6% yesno-rev same top passage as noul on 90.2% ab same top passage as noul on 90.0% ab-rev same top passage as noul on 88.9% xy same top passage as noul on 88.7% xy-rev same top passage as noul on 84.1% relevant same top passage as noul on 87.3% relevant-rev same top passage as noul on 86.7% choice-avg same top passage as noul on 90.2% yesno vs its duplicate: 95.4% cost of Part A: $2.57 for 1617 questions ============================================================================================================== PART B. Pick-one mode with different passage ids 1617 questions with every run (check: picks that are not the highest-probability passage: 3) run top1 nDCG10 ECE said none >=0.9 share >=0.9 right >=0.95 share >=0.95 right p01 ids, published order 76.0% 0.733 0.052 10.7% 37.4% 93.2% (90.9% to 95.0%) 28.8% 95.5% p01 ids, asked again 75.6% 0.732 0.049 10.5% 38.2% 91.9% (89.5% to 93.8%) 28.8% 95.7% p01 ids, shuffled 75.1% 0.731 0.037 10.6% 36.9% 93.0% (90.6% to 94.8%) 28.8% 96.1% numbers, published order 76.3% 0.734 0.051 10.3% 36.7% 92.4% (90.0% to 94.3%) 28.3% 95.6% numbers, shuffled 75.5% 0.731 0.043 10.4% 37.4% 92.5% (90.2% to 94.4%) 29.1% 96.4% letters, published order 76.2% 0.733 0.052 9.8% 37.9% 93.0% (90.7% to 94.8%) 29.2% 95.6% letters, shuffled 75.1% 0.730 0.035 10.2% 37.4% 93.5% (91.3% to 95.2%) 29.9% 95.7% random ids, published order 76.2% 0.735 0.049 10.1% 38.3% 92.9% (90.6% to 94.7%) 29.1% 96.4% random ids, shuffled 76.1% 0.733 0.044 10.2% 36.9% 94.8% (92.7% to 96.3%) 29.4% 96.6% Top passage changed compared with p01 ids in the published order: p01 ids, asked again 2.9% (2.2% to 3.8%) p01 ids, shuffled 20.5% (18.6% to 22.6%) numbers, published order 4.3% (3.4% to 5.4%) numbers, shuffled 20.7% (18.8% to 22.7%) letters, published order 4.7% (3.8% to 5.8%) letters, shuffled 20.2% (18.3% to 22.2%) random ids, published order 5.5% (4.5% to 6.7%) random ids, shuffled 19.0% (17.1% to 21.0%) Top passage changed between the published and shuffled order, same ids: p01 ids 20.5% numbers 20.3% letters 20.0% random ids 18.5% By the confidence of the published run: do the 4 id schemes (published order) agree? confidence questions all agree asked-again agrees conf spread 0.00 to 0.50 345 68.4% 89.0% 0.091 0.50 to 0.70 327 91.4% 98.2% 0.109 0.70 to 0.90 341 98.8% 99.4% 0.083 0.90 to 0.95 138 98.6% 99.3% 0.042 0.95 to 0.99 214 100.0% 100.0% 0.019 0.99 to 1.00 252 100.0% 100.0% 0.004 cost of Part B (all rows written, including the test rows): $3.31