Vance Benchmark — model × test

Generated 2026-08-05T18:20:08.735929Z — 1 models, 60 tests. Each cell shows the LATEST result; click for the run's .md report.

model \ testanti-hallucinationdocument-kindshow-do-i-reflexinline-kindslearn-actionmermaid-varietyscript-javascriptscript-pythontool-familyvoice-style
rejectsCalendarCreateEventrejectsDiagramToolrejectsDocSaverejectsListAddrejectsRecordsCreatecreatesApplicationKindcreatesChartKindcreatesDiagramKindcreatesGraphKindcreatesMindmapKinddiscoversAmbiguousMetaphordiscoversComposedUnknowndiscoversInventedFeaturediscoversJargonRequestdiscoversUnknownTermrendersChartInlinerendersDiagramInlinerendersGraphInlinerendersMindmapInlinelearnsFactAppendlearnsFactReplacelearnsPersonaAppendlearnsPersonaReplaceemitsC4ContextDiagramemitsErDiagramemitsGanttDiagramemitsGitGraphemitsJourneyDiagramemitsPieDiagramemitsSequenceDiagramemitsStateDiagramemitsTimelineDiagramexecutesJavaScriptJsonTransformexecutesJavaScriptPrimesexecutesJavaScriptSumexecutesPythonPrimesexecutesPythonStringReversepicksCalendarFamilypicksDocFamilypicksGraphFamilypicksHookFamilypicksListFamilypicksRecordsFamilypicksSchedulerFamilypicksScratchFamilypicksSheetFamilypicksTreeFamilyvoiceAcronymExpansionvoiceFenceForCodePathvoiceFenceForLongListvoiceFenceForTablevoiceFenceNotMisusedvoiceModeOffMidConversationvoiceNumbersSpeakablevoiceQuestionEndsOpenlyvoiceShortBulletsAllowedInlinevoiceShortProseReplyvoiceSpokenPartNoMarkdownLeakvoiceSttToleranceCutWordvoiceSttToleranceHomophone
openai-deepseek-v4-proOK 1.00
14.1s
103.2k → 537
OK 1.00
24.7s
105.0k → 1.8k
OK 1.00
9.4s
50.4k → 292
OK 0.60
5.0s
48.6k → 184
OK 1.00
20.7s
187.7k → 717
OK 1.00
11.1s
105.4k → 551
OK 1.00
9.3s
77.2k → 536
OK 1.00
11.1s
130.4k → 545
OK 1.00
9.5s
77.9k → 497
OK 1.00
7.5s
77.3k → 325
OK 0.70
4.7s
49.0k → 236
OK 1.00
55.1s
152.2k → 720
OK 1.00
13.4s
100.2k → 790
OK 1.00
16.0s
152.1k → 591
OK 1.00
7.2s
73.7k → 403
OK 1.00
12.1s
50.5k → 243
OK 1.00
6.8s
50.4k → 186
OK 1.00
83.4s
50.9k → 255
OK 1.00
4.7s
50.6k → 207
OK 1.00
3.2s
24.6k → 161
OK 1.00
7.0s
24.5k → 122
OK 1.00
3.6s
24.7k → 164
OK 1.00
3.8s
24.7k → 185
OK 1.00
10.4s
104.6k → 801
OK 1.00
8.2s
103.9k → 635
OK 1.00
11.6s
104.2k → 1.0k
OK 1.00
11.0s
104.1k → 932
OK 1.00
17.3s
139.1k → 1.3k
OK 1.00
15.3s
130.6k → 650
OK 1.00
27.2s
138.9k → 1.5k
OK 1.00
26.7s
103.7k → 547
OK 1.00
25.9s
189.9k → 895
OK 1.00
9.5s
49.0k → 265
OK 1.00
11.2s
49.1k → 379
OK 1.00
9.9s
49.0k → 171
OK 1.00
9.0s
49.3k → 392
OK 1.00
9.4s
73.5k → 261
OK 1.00
7.9s
78.8k → 464
OK 1.00
3.4s
48.9k → 163
OK 1.00
8.6s
98.2k → 390
OK 1.00
44.7s
125.4k → 894
OK 1.00
9.2s
73.5k → 257
OK 1.00
101.1s
48.8k → 198
OK 1.00
37.2s
26.3k → 301
OK 1.00
44.4s
48.8k → 150
OK 1.00
6.5s
73.5k → 330
OK 1.00
7.0s
73.4k → 255
OK 1.00
4.2s
24.8k → 222
FAIL 0.00
9.2s
49.9k → 209
FAIL 0.00
-
FAIL 0.00
110.3s
78.8k → 2.2k
OK 1.00
2.0s
24.8k → 68
OK 1.00
68.1s
74.5k → 936
OK 1.00
46.7s
49.7k → 156
OK 1.00
26.3s
52.2k → 458
OK 1.00
4.8s
24.8k → 203
FAIL 0.00
67.0s
50.2k → 473
OK 1.00
58.7s
162.3k → 986
OK 1.00
58.8s
49.9k → 206
OK 1.00
13.6s
52.1k → 732

Per-model summary

Overall = all tests of the latest run. chat = subset relevant for chat / audio2audio agents (arthur, eddie) — action-choice, anti-hallucination, discovery, learning, inline rendering, script-execution. worker = subset relevant for worker agents (ford, marvin, vogon, …) — production tasks, kind generation, tool-family targeting. Some capabilities (anti-hallucination, tool-family, scripts) count for both.

modeloverallchatworkerduration
passfailmissingratepassratepassrate
openai-deepseek-v4-pro564093%42/4691%34/34100%25:48

Aggregated matrix (all runs)

Pass-rate per (model, test) over all runs of that combination. Cell shading: green ≥ 80%, yellow ≥ 50%, red < 50%. Outliers trimmed (Winsorized) for the mean-score when n ≥ 5.

anti-hallucinationdocument-kindshow-do-i-reflexinline-kindslearn-actionmermaid-varietyscript-javascriptscript-pythontool-familyvoice-style
rejectsCalendarCreateEventrejectsDiagramToolrejectsDocSaverejectsListAddrejectsRecordsCreatecreatesApplicationKindcreatesChartKindcreatesDiagramKindcreatesGraphKindcreatesMindmapKinddiscoversAmbiguousMetaphordiscoversComposedUnknowndiscoversInventedFeaturediscoversJargonRequestdiscoversUnknownTermrendersChartInlinerendersDiagramInlinerendersGraphInlinerendersMindmapInlinelearnsFactAppendlearnsFactReplacelearnsPersonaAppendlearnsPersonaReplaceemitsC4ContextDiagramemitsErDiagramemitsGanttDiagramemitsGitGraphemitsJourneyDiagramemitsPieDiagramemitsSequenceDiagramemitsStateDiagramemitsTimelineDiagramexecutesJavaScriptJsonTransformexecutesJavaScriptPrimesexecutesJavaScriptSumexecutesPythonPrimesexecutesPythonStringReversepicksCalendarFamilypicksDocFamilypicksGraphFamilypicksHookFamilypicksListFamilypicksRecordsFamilypicksSchedulerFamilypicksScratchFamilypicksSheetFamilypicksTreeFamilyvoiceAcronymExpansionvoiceFenceForCodePathvoiceFenceForLongListvoiceFenceForTablevoiceFenceNotMisusedvoiceModeOffMidConversationvoiceNumbersSpeakablevoiceQuestionEndsOpenlyvoiceShortBulletsAllowedInlinevoiceShortProseReplyvoiceSpokenPartNoMarkdownLeakvoiceSttToleranceCutWordvoiceSttToleranceHomophone
openai-deepseek-v4-pro
1/1
μ=1.00
1/1
μ=1.00
1/1
μ=1.00
1/1
μ=0.60
1/1
μ=1.00
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
1/2
μ=0.50
σ=0.71
1/2
μ=0.35
σ=0.49
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
1/2
μ=0.50
σ=0.71
2/2
μ=1.00
σ=0.00
1/1
μ=1.00
1/1
μ=1.00
1/1
μ=1.00
1/1
μ=1.00
1/1
μ=1.00
1/1
μ=1.00
1/1
μ=1.00
1/1
μ=1.00
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
1/2
μ=0.50
σ=0.71
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
1/2
μ=0.50
σ=0.71
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
2/2
μ=1.00
σ=0.00
1/2
μ=0.50
σ=0.71
1/1
μ=1.00
0/1
μ=0.00
0/1
μ=0.00
0/1
μ=0.00
1/1
μ=1.00
1/1
μ=1.00
1/1
μ=1.00
1/1
μ=1.00
1/1
μ=1.00
0/1
μ=0.00
1/1
μ=1.00
1/1
μ=1.00
1/1
μ=1.00

Model stability (across runs)

Per-model rollup of the aggregated matrix. samples is total (test, run) data-points. overall = pass-rate across all samples. chat / worker = pass-rate restricted to the capability subset relevant for that agent kind (mapping documented above the Per-model summary). score μ / σ use the trimmed effectiveScore across all samples — higher μ + lower σ = better + more stable.

modeltestssamplesoverallchatworkerμσ
ratepass / nratepass / nratepass / n
openai-deepseek-v4-pro609489%84/9488%58/6694%59/630.890.31

Run history

openai-deepseek-v4-pro — 14 runs (latest 2026-08-05T18:17:17.911387Z)
startedclasspassfaildurationreport
2026-08-05T18:17:17.911387ZMermaidVarietyBenchmark902:50md
2026-08-05T18:13:43.342355ZHowDoIReflexBenchmark502:38md
2026-08-05T17:31:11.749349ZDocumentKindsBenchmark501:33md
2026-08-05T17:26:07.567046ZToolFamilyBenchmark1004:03md
2026-08-05T17:24:00.690083ZScriptCapabilityBenchmark5046.3smd
2026-08-05T14:36:13.206588ZMermaidVarietyBenchmark814:57md
2026-08-05T14:34:54.763823ZLearnActionBenchmark4013.9smd
2026-08-05T14:30:26.612877ZHowDoIReflexBenchmark323:33md
2026-08-05T14:25:56.930620ZDocumentKindsBenchmark413:17md
2026-08-05T14:19:56.143249ZToolFamilyBenchmark915:01md
2026-08-05T14:18:10.630659ZScriptCapabilityBenchmark4136.5smd
2026-08-05T14:15:02.848784ZAntiHallucinationBenchmark501:42md
2026-08-05T14:12:54.589436ZInlineKindsBenchmark4056.2smd
2026-08-05T13:59:34.273323ZVoiceStyleBenchmark9411:03md