Graph Learning · LLM × Graph · Multi-Agent · Science
Showing 9 papers for 2026-09-24
Nothing cleared the bar today. Papers read in full today: 2.
The central idea is genuinely load-bearing: isolate graph access from capacity/training confounds, then audit whether the benchmark is mostly measuring a service-name prior rather than RCA reasoning. The full text supports the negative claim well—matched controls show only a 0.003 in-distribution gap, the prior-only ranker gets Avg@5 0.488, and the graph residual adds little beyond the prior—so the main result is a solid demystification rather than a new method win.
The main limitation is that the strongest positive claims are exploratory and underpowered: only three systems for transfer, one architecture family, and the PSC-GRCA decomposition is scored from one trained model rather than showing what training without the prior would do.
The Tasteful Agent defines the notion of 'taste' as the ability to make good long-horizon decisions that shape the outcome of extended tasks. It proposes Taste-Bench, a benchmark of taste questions designed to probe this decision quality beyond end-task success. The work emphasizes that measuring taste enables diagnosing weaknesses in long-horizon planning and guiding targeted improvements in agents' decision strategies.
RULER addresses the lack of a faithful signal for SVG generation by moving away from scalar metrics toward rubric-based scoring. It uses a vision-language judge guided by a multi-axis rubric, and shows that this rubric-based signal correlates with human judgments far better than scalar metrics, reducing reward-hacking risk and improving policy optimization for SVG generation.
GAE introduces a geometry-native latent space as a shared foundation for perception and generation, addressing the mismatch where generators produce photorealistic frames without 3D consistency. It argues that perception recovers geometry in a rich, cross-view space, and reparameterizes a geometry foundation model's features into a compact latent space for generation, enabling 3D-consistent world generation from a single representation.
All-in-One Multilingual Scene Text Recognition proposes a script-aware mixture-of-experts approach to build a single model that handles multiple languages and scripts. It introduces TextMuSS-10M, a large multilingual dataset, and demonstrates that the unified model achieves strong multilingual accuracy with lower cost than per-language systems and greater efficiency than large vision-language models.
Ovis-Embedding presents a universal omni-modal embedding family built on a shared multimodal backbone for text, image, video, and audio. It advances three directions: native omni-modal initialization using a pretrained Qwen-omni backbone with low-rank contrastive adaptation; data-centric omni-modal training to curate high-quality multimodal data; and broad modality coverage to enable versatile cross-modal tasks.
This work shows that errors in node classification with GNNs can stem from two separate causes: poor mixture weights over local messages or mispositioned reachable logits. They formalize this with an exact-mass linear program and propose two learned post-hoc repairs; reweighting yields translations in logit space that are only realizable within a message-induced displacement set. Across eight datasets and eight backbones, these repairs yield meaningful accuracy gains and more reliable predictions.
We propose SCoGL, a spectral connectivity-regularized graph learning framework for scarce data. SCoGL augments a Laplacian-constrained graphical Lasso objective over an adjacency W with Laplacian spectral priors designed to promote global connectivity while preserving sparsity. In scarce-data regimes, these priors improve graph reconstruction and downstream task performance.
We propose QUARTET, a quad-branch cross-attention architecture with random-walk traces to enhance transformers on relational graphs. It directly tackles RelGT's limitations: loosely connected subgraphs caused by local samplers and a single seed-feature-based global memory that misses macro-level dynamics. By enriching both local message passing and global context, QUARTET improves relational graph transformer performance on benchmarks.
We propose resistance-curvature-guided subgraph sampling (ERC-LG) to preserve geometric roles of edges while scaling training for large graphs. ERC-LG uses Johnson-Lindenstrauss projections and regularized multi-GPU batched conjugate gradient solvers to estimate curvature without computing Laplacian pseudoinverse or storing full embeddings. The resulting curvature-informed sampling probabilities improve efficiency and maintain or improve accuracy in large-scale GNN training.
Dual-Hypergraph Indexing (DHI) is proposed to bridge knowledge islands for multi-hop reasoning in retrieval-augmented generation. It introduces a hierarchical dual-hypergraph representation that elevates discrete facts into structured analytical insights, enabling better multi-hop causal inference and narrative synthesis.
LEGO is a dual-module framework that synergizes expert GraphRAG and expert Chain-of-Thought to improve legal reasoning with large language models. It addresses two bottlenecks: normative relations among legal provisions are underutilized by current retrieval-based methods, and vanilla CoT prompts may generate plausible but normatively flawed rationales. By aligning retrieval with normative reasoning in the graph and guiding CoT with normative structure, LEGO enhances legal accuracy and transparency.
This work studies curriculum learning for GNN-based reinforcement learning in the job shop scheduling problem. Large instances make training expensive, so they compare curricula that progressively increase problem size against single-size training, evaluating generalization across instance sizes. Results show that curriculum-based training can improve efficiency and generalization relative to one-shot training.
We present a benchmarking suite for Automated Knowledge Graph Construction from Semi-Structured Data. While KGs are increasingly used, there is a lack of comprehensive benchmarks for semi-structured inputs, as opposed to text-only methods. The benchmark defines datasets, tasks, evaluation metrics, and baseline methods, enabling systematic evaluation and progress in semi-structured KG construction.