Highlights · Graph cognition · Graph + text · Sequential decisions · Full tables · Cite
How well do fast decision models understand graphs—and turn choices into good solutions? GraphDecide compares Jev, related native selectors and language models through three complementary tests. The benchmark is model-independent; a common choice interface does not imply a common architecture or reasoning budget.
Results: manuscript v33 · October 2026. All figures below use the primary v33 panels, not the older supporting experiments. GPT‑5.4 uses no reasoning; GPT‑6‑Astra is a reasoning-enabled reference. There is no cross-task overall rank.
| Graph cognition | Graph + text | Sequential decisions |
|---|---|---|
| Local recognition ≠ structural mastery. | More evidence ≠ better decisions. | Completion ≠ solution quality. |
| Jev reaches 100% adjacency, but 67.80% degree and 61.25% articulation. | Jev's explicit-edge gain is +1 pp on arXiv and +3 pp on Prime; both paired intervals include zero. | Proposals reduce Jev's TSP gap from 59.66% to 25.87%, yet a fixed proposal rule reaches 15.58%. |
The value of GraphDecide is the profile, not a single leaderboard score: structural correctness, information utility and construction quality answer different questions.
Six tasks · four real networks · 8–12 vertices · 800 queries per configuration.
Full-size figure · Exact values · GPT-5.4 query-level CSV
- Jev is strong but uneven. It matches the large-model baselines on adjacency, but the gap on degree and articulation is substantial.
- Scale weighting matters. GPT‑5.4's degree accuracy is 91.4154% when sizes receive equal weight, versus 91.0000% when queries are pooled.
- The strongest reference is not compute-matched. Astra reaches 100% across the six tasks with reasoning enabled.
Metric and dataset notes
Sources: Facebook, ca-GrQc, Power Grid and Human PPI, 200 queries each. There are 400 degree questions and 80 for each other task. The primary task score averages accuracy equally across sizes 8–12; illegal outputs count as incorrect. Labels are computed on exactly the supplied graph, including isolates.
| GPT-5.4 aggregation | Accuracy |
|---|---|
| Query weighted, 733/800 correct | 91.6250% |
| Task macro, equal weight per pooled task | 92.0417% |
| Task-size macro, equal weight per task and size | 92.1109% |
All 800 GPT-5.4 final outputs are legal. The heatmap follows v33 Appendix C; each task is a separate comparison, not part of a blended score.
Two datasets · five matched inputs · 100 test queries per dataset and condition.
Full-size figure · arXiv table · Prime table · GPT-5.4 paired CSV
T = text · G = relations · TG = joint input · BAG = TG without explicit edges · A/A* = comparator without additional entity text or edges.
| Jev paired contrast | ogbn-arxiv | STaRK-Prime |
|---|---|---|
| TG − BAG: explicit edges, same accompanying text | +1 pp [−4, +7] | +3 pp [−6, +12] |
| G − A/A*: relations without additional entity text | +10 pp [+3, +18] | +5 pp [−5, +14] |
Intervals are paired, pointwise 95% bootstrap intervals—not recoverable from aggregate correct counts. Relational input helps Jev over arXiv anchors alone, but adding edges to text-rich input has no clear positive effect. High TG accuracy and a positive TG−BAG gain are different claims.
What these comparisons do—and do not—measure
- arXiv has forty classes and shared training-label anchors. Context identities are graph-selected even in T, so it is not a graph-independent baseline. TG−T adds context text as well as edges.
- Prime is selection within forty gold-containing candidates, not end-to-end retrieval. A missing native answer is inserted deterministically into BM25 top-40; candidates are shuffled without answer markers. Candidate Hit@40 is 100% by design; Recall@40 is 77.7241%.
Umeans unsupported, not zero accuracy. GPT‑5.4 retains two illegal Prime/TG outcomes as wrong: 603/1,000 correct, 998/1,000 legal.- arXiv macro-F1 averages over the fixed forty classes, including zero F1 for classes absent from both truth and predictions. It cannot be inferred from binary correctness alone.
Four optimization tasks · ten graphs each · direct actions versus heuristic proposals.
Full-size figure · Exact gaps and coverage · GPT-5.4 trajectory CSV
A — Direct construction: choose among all legal next actions.
C — Heuristic-proposal selection: choose among distinct next actions proposed by four fixed rules. A singleton is forced, without a model call.
| Jev task | Direct gap | Proposal gap | Fixed-rule comparison | Complete A / C |
|---|---|---|---|---|
| TSP | 59.66% | 25.87% | return_reserve: 15.58% |
10/10 · 10/10 |
| Signed MaxCut | 99.96% | 9.05% | greedy_completion: 8.87% |
10/10 · 10/10 |
| Deterministic LT | 52.84% | 0.50% | Partial cohort—compare on matched graphs | 8/10 · 8/10 |
| Binary modularity | 0.4784 Q | 0.0555 Q | greedy_completion: 0.0361 Q |
10/10 · 10/10 |
Lower is better. Strong improvement over direct construction does not establish improvement over the proposal rules themselves. Proposals also do not help every configuration: GPT‑5.4's TSP gap increases from 29.49% to 32.24%; Astra's direct policy is better than its proposal policy on all four tasks.
Coverage, references and fair comparisons
- Across fourteen configurations, 951/1,120 trajectories complete, 165 are unsupported and four fail. Jev completes 76/80; its four failures are context-capacity HTTP errors on the two 150-node LT graphs, in both modes.
- Means use completed trajectories. Compare modes and controls on the same completed-instance subset; a partial result is not rankable against a full-ten-graph result. Forced-only completion does not establish interface support.
- TSP uses published optima, LT exhaustive two-seed optima, signed MaxCut historical 2016 BKS, and binary modularity numerical MILP-certified optima. TSP/LT/MaxCut gaps are percentages; modularity gaps are absolute Q.
- Signed gaps are not clamped. Non-GPT modularity objectives were rounded to four decimals before rebasing; additional printed digits are not extra precision.
- GPT‑5.4's 80 rows are final selections after invalid-only retries, not a single-attempt completion-rate claim. A legal answer was retained even when wrong. Main CSV call counts refer only to the selected trajectory; the separate retry table retains all extra costs.
| Resource | What it contains |
|---|---|
| Complete result tables | All fourteen configurations, taskwise scores, gaps and completion counts |
| Figure source data | v33 manuscript-transcribed values; all charts and full tables share this source |
| GPT-5.4 final-result bundle | 800 structural rows, 1,000 paired graph-text rows, 80 trajectories, checksums |
| Reproduction and model settings | Offline reporting, configuration, tests and fresh-clone limitations |
| Historical result assets | Separate supporting cohorts; not pooled with the v33 primary panels |
Readouts differ. Native selection, Qwen constrained generation and
candidate-token scoring retain their own support limits. GPT‑5.4 uses
reasoning_effort="none" with verified zero reasoning tokens; Astra uses
observed medium reasoning. Copilot's tool enum is not verified strict
decoding. These comparisons do not establish a matched-hardware speed ranking.
If you use the benchmark, code or results, please cite GraphDecide and identify the version or commit used:
@misc{yang2026graphdecide,
title = {GraphDecide: Benchmarking System One Models on Graph Tasks},
author = {Yang, Xianliang and Zhang, Yapu and Zhao, Li},
year = {2026},
howpublished = {Manuscript},
note = {Version 33},
url = {https://github.com/VictorYXL/JevGraphBench}
}Machine-readable citation. This is a manuscript citation; no publication venue, DOI or arXiv identifier is asserted. If reusing data, also cite the original sources: SNAP, Power Grid, BioSNAP, OGB, STaRK, TSPLIB and OPTSICOM. Source-specific provenance and citation requirements are described in the data documentation and manuscript.