Skip to content

About

Benchmarking Jev’s Capabilities on Graph Tasks

Resources

Stars

0 stars

Watchers

0 watching

Forks

 
 

Latest commit

 

History

12 Commits

Folders and files

Repository files navigation

GraphDecide: Benchmarking System One Models on Graph Tasks. Fourteen model configurations; 800 structural queries, 1,000 graph-text decisions and 80 construction trajectories per configuration.

Highlights  ·  Graph cognition  ·  Graph + text  ·  Sequential decisions  ·  Full tables  ·  Cite

How well do fast decision models understand graphs—and turn choices into good solutions? GraphDecide compares Jev, related native selectors and language models through three complementary tests. The benchmark is model-independent; a common choice interface does not imply a common architecture or reasoning budget.

Results: manuscript v33 · October 2026. All figures below use the primary v33 panels, not the older supporting experiments. GPT‑5.4 uses no reasoning; GPT‑6‑Astra is a reasoning-enabled reference. There is no cross-task overall rank.

Key findings

Graph cognition Graph + text Sequential decisions
Local recognition ≠ structural mastery. More evidence ≠ better decisions. Completion ≠ solution quality.
Jev reaches 100% adjacency, but 67.80% degree and 61.25% articulation. Jev's explicit-edge gain is +1 pp on arXiv and +3 pp on Prime; both paired intervals include zero. Proposals reduce Jev's TSP gap from 59.66% to 25.87%, yet a fixed proposal rule reaches 15.58%.

The value of GraphDecide is the profile, not a single leaderboard score: structural correctness, information utility and construction quality answer different questions.

RQ1 — Graph cognition

Six tasks · four real networks · 8–12 vertices · 800 queries per configuration.

RQ1: size-equal accuracy for fourteen model configurations across adjacency, degree, cycle, connectivity, distance and articulation. Darker cells indicate higher accuracy.

Full-size figure · Exact values · GPT-5.4 query-level CSV

  • Jev is strong but uneven. It matches the large-model baselines on adjacency, but the gap on degree and articulation is substantial.
  • Scale weighting matters. GPT‑5.4's degree accuracy is 91.4154% when sizes receive equal weight, versus 91.0000% when queries are pooled.
  • The strongest reference is not compute-matched. Astra reaches 100% across the six tasks with reasoning enabled.
Metric and dataset notes

Sources: Facebook, ca-GrQc, Power Grid and Human PPI, 200 queries each. There are 400 degree questions and 80 for each other task. The primary task score averages accuracy equally across sizes 8–12; illegal outputs count as incorrect. Labels are computed on exactly the supplied graph, including isolates.

GPT-5.4 aggregation Accuracy
Query weighted, 733/800 correct 91.6250%
Task macro, equal weight per pooled task 92.0417%
Task-size macro, equal weight per task and size 92.1109%

All 800 GPT-5.4 final outputs are legal. The heatmap follows v33 Appendix C; each task is a separate comparison, not part of a blended score.

RQ2 — Graph and text

Two datasets · five matched inputs · 100 test queries per dataset and condition.

RQ2: arXiv accuracy and STaRK-Prime Hit@1 for all fourteen configurations under T, G, TG, BAG and A or A-star. Unsupported interfaces are marked U, not zero.

Full-size figure · arXiv table · Prime table · GPT-5.4 paired CSV

T = text · G = relations · TG = joint input · BAG = TG without explicit edges · A/A* = comparator without additional entity text or edges.

Do explicit edges add value?

Jev paired contrast ogbn-arxiv STaRK-Prime
TG − BAG: explicit edges, same accompanying text +1 pp [−4, +7] +3 pp [−6, +12]
G − A/A*: relations without additional entity text +10 pp [+3, +18] +5 pp [−5, +14]

Intervals are paired, pointwise 95% bootstrap intervals—not recoverable from aggregate correct counts. Relational input helps Jev over arXiv anchors alone, but adding edges to text-rich input has no clear positive effect. High TG accuracy and a positive TG−BAG gain are different claims.

What these comparisons do—and do not—measure
  • arXiv has forty classes and shared training-label anchors. Context identities are graph-selected even in T, so it is not a graph-independent baseline. TG−T adds context text as well as edges.
  • Prime is selection within forty gold-containing candidates, not end-to-end retrieval. A missing native answer is inserted deterministically into BM25 top-40; candidates are shuffled without answer markers. Candidate Hit@40 is 100% by design; Recall@40 is 77.7241%.
  • U means unsupported, not zero accuracy. GPT‑5.4 retains two illegal Prime/TG outcomes as wrong: 603/1,000 correct, 998/1,000 legal.
  • arXiv macro-F1 averages over the fixed forty classes, including zero F1 for classes absent from both truth and predictions. It cannot be inferred from binary correctness alone.

RQ3 — Sequential decisions

Four optimization tasks · ten graphs each · direct actions versus heuristic proposals.

RQ3: direct and heuristic-proposal gaps for every configuration on TSP, signed MaxCut, deterministic LT and binary modularity. Lower is better. Dashed bars and bracketed counts identify partial coverage; each task has a separate axis.

Full-size figure · Exact gaps and coverage · GPT-5.4 trajectory CSV

A — Direct construction: choose among all legal next actions.

C — Heuristic-proposal selection: choose among distinct next actions proposed by four fixed rules. A singleton is forced, without a model call.

Proposals help—but how much is the model contributing?

Jev task Direct gap Proposal gap Fixed-rule comparison Complete A / C
TSP 59.66% 25.87% return_reserve: 15.58% 10/10 · 10/10
Signed MaxCut 99.96% 9.05% greedy_completion: 8.87% 10/10 · 10/10
Deterministic LT 52.84% 0.50% Partial cohort—compare on matched graphs 8/10 · 8/10
Binary modularity 0.4784 Q 0.0555 Q greedy_completion: 0.0361 Q 10/10 · 10/10

Lower is better. Strong improvement over direct construction does not establish improvement over the proposal rules themselves. Proposals also do not help every configuration: GPT‑5.4's TSP gap increases from 29.49% to 32.24%; Astra's direct policy is better than its proposal policy on all four tasks.

Coverage, references and fair comparisons
  • Across fourteen configurations, 951/1,120 trajectories complete, 165 are unsupported and four fail. Jev completes 76/80; its four failures are context-capacity HTTP errors on the two 150-node LT graphs, in both modes.
  • Means use completed trajectories. Compare modes and controls on the same completed-instance subset; a partial result is not rankable against a full-ten-graph result. Forced-only completion does not establish interface support.
  • TSP uses published optima, LT exhaustive two-seed optima, signed MaxCut historical 2016 BKS, and binary modularity numerical MILP-certified optima. TSP/LT/MaxCut gaps are percentages; modularity gaps are absolute Q.
  • Signed gaps are not clamped. Non-GPT modularity objectives were rounded to four decimals before rebasing; additional printed digits are not extra precision.
  • GPT‑5.4's 80 rows are final selections after invalid-only retries, not a single-attempt completion-rate claim. A legal answer was retained even when wrong. Main CSV call counts refer only to the selected trajectory; the separate retry table retains all extra costs.

Results, data and scope

Resource What it contains
Complete result tables All fourteen configurations, taskwise scores, gaps and completion counts
Figure source data v33 manuscript-transcribed values; all charts and full tables share this source
GPT-5.4 final-result bundle 800 structural rows, 1,000 paired graph-text rows, 80 trajectories, checksums
Reproduction and model settings Offline reporting, configuration, tests and fresh-clone limitations
Historical result assets Separate supporting cohorts; not pooled with the v33 primary panels

Readouts differ. Native selection, Qwen constrained generation and candidate-token scoring retain their own support limits. GPT‑5.4 uses reasoning_effort="none" with verified zero reasoning tokens; Astra uses observed medium reasoning. Copilot's tool enum is not verified strict decoding. These comparisons do not establish a matched-hardware speed ranking.

Citation

If you use the benchmark, code or results, please cite GraphDecide and identify the version or commit used:

@misc{yang2026graphdecide,
  title        = {GraphDecide: Benchmarking System One Models on Graph Tasks},
  author       = {Yang, Xianliang and Zhang, Yapu and Zhao, Li},
  year         = {2026},
  howpublished = {Manuscript},
  note         = {Version 33},
  url          = {https://github.com/VictorYXL/JevGraphBench}
}

Machine-readable citation. This is a manuscript citation; no publication venue, DOI or arXiv identifier is asserted. If reusing data, also cite the original sources: SNAP, Power Grid, BioSNAP, OGB, STaRK, TSPLIB and OPTSICOM. Source-specific provenance and citation requirements are described in the data documentation and manuscript.

About

Benchmarking Jev’s Capabilities on Graph Tasks

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages