Graph Learning · LLM × Graph · Multi-Agent · Science
Showing 71 papers for 2026-09-29
Nothing cleared the bar today. Papers read in full today: 24.
The central idea is load-bearing: conventional Recall/HR/NDCG can hide the fact that sign-aware recommenders still rank disliked items highly, and the paper supports that with V-AUC, linear probes, and the large re-ranking under Signed Recall/HR/NDCG. The main concrete finding is that several strong sign-aware models sit near chance on positive-vs-negative separation, and that a simple auxiliary sign loss plus valence adjustment improves signed metrics on most datasets, so the metric is not just cosmetic. However, the paper does not propose a stronger core recommender; it mainly shows that the community has been measuring the wrong thing.
The signed metric is a plausible fix for evaluation, but the proof-of-concept training tweak is lightweight and the claim is narrower than the title sounds: it is mainly a measurement reform, not a new model or theory of signed recommendation.
The central idea is that reward should be projected through the forward corruption process, while a second controller reallocates mass over graph sizes so the model can move the terminal distribution instead of only refining structures at a fixed size. The paper does support this with a nontrivial theorem and a useful ablation: GraphFDM w/o SDC is already strong, but SDC further improves MAE and validity, which suggests the size-marginal story is real rather than decorative. Still, the empirical gain over the strongest post-training baselines is incremental, and the benchmark story is narrower than the theory sounds.
The main limitation is that the headline improvement is still mostly a stronger variant of reward-based post-training on two molecular benchmarks, so the paper changes practice less than the theory suggests.
The core idea is load-bearing: turn quadrilateral block decomposition into an MDP whose reward is the Gauss–Bonnet lower bound, then use backward-generated optimal trajectories for cloning because random exploration almost never reaches the certificate. The full text does support the claim with strong results against Gmsh and an ablation showing scratch PPO essentially fails, so this is more than decoration on a standard RL loop. Still, the novelty is mainly in problem formulation plus training recipe for a specialized geometry task, not a new general graph-learning principle.
The strongest limitation is that the solved criterion depends on a smoother in the loop and the method needs test-time repair for curved boundaries, so the certificate is less self-contained than the headline suggests.
The load-bearing idea is the MDL objective that turns hyperedge recovery into “compression gain minus specification cost,” yielding a local searchable criterion and a detectability threshold. The text does support that claim: it derives the objective, shows the boundary matches planted-dense-subgraph recovery scaling in the single-interaction limit, and then the synthetic experiments show the λ_h>1 boundary strongly predicts which planted interactions are recovered. The delivery is decent rather than transformative: the empirical win is mostly against a Bayesian clique-based baseline and a motif-cover method that is not competitive at scale, and the algorithm is still a heuristic local search rather than a guaranteed optimizer. So the paper is valuable as a scalable, theory-backed reconstruction method, but the evidence says “better and more principled hypergraph reconstruction,” not a new subfield-defining capability.
The main limitation is that the headline scalability and recovery claims rest on a heuristic local search and on baselines that are either model-mismatched or weak at large scale, so the strongest guarantees are only for the single-interaction asymptotic analysis, not the full multi-interaction algorithm.
The core idea is clear and load-bearing: convert hindsight credit into a scalar success potential on states, then estimate it on a pooled transition graph with a contraction argument, avoiding a learned hindsight model. The paper does test that claim: on ALFWorld, WebShop, and Sokoban it beats GRPO, GiGPO, HCAPO, and GraphGPO, and the ablation with power-mean aggregation shows the expectation backup matters rather than just any graph backup. The honest read is that this is a solid refinement of step-level credit assignment, not a paradigm shift.
The guarantee and closed form rely on deterministic transitions and terminal binary rewards, so the claim is narrower than the LLM-agent framing suggests.
The load-bearing idea is real: using initialization-dependent attractors to represent a set of valid graph solutions instead of forcing one arbitrary target. The full text does support that story with a theorem separating multi-attractor from contractive dynamics and with experiments showing multiple good solutions on Ising, protein modularity, and reaction-network tasks, but the empirical section is fairly small and mostly compares against unique-equilibrium and single-target controls rather than strongest modern search/generative baselines under equal budgets.
The strongest limitation is that the theory is representational and asymptotic, while the experiments do not certify exact solution coverage or isolate recurrence from budget differences against modern generative solvers.
The core idea is a taxonomy of long-range failure modes: architectural support limit, approximation gap, alignment gap, and execution gap, plus a distance-based influence profile/R_q diagnostic. The full text does support the claim that these gaps are real and can be separated in controlled settings, but the empirical contribution is mostly diagnostic rather than transformative: it shows that several GNNs with similar nominal reach behave very differently, and that low average error can hide catastrophic loss of distant tail behavior.
The main limitation is that many of the insights are assembled from existing sensitivity, implicit-operator, and benchmark ideas; the paper sharpens the story, but it does not convincingly establish a new method that practitioners would adopt instead of existing long-range models.
FuseReg regularizes layer fusion in representation autoencoders to mitigate the reconstruction–generation gap. By aligning which encoder layers form the shared latent space for the generator and the pixel decoder, FuseReg decouples the conflicting information carried by shallow versus deep layers and reduces the trade‑off between preserving pixel detail and achieving strong generation metrics. The result is improved end-to-end reconstruction and synthesis quality in RAEs.
Disaggregated Quantization (DQ) specializes quantization for both prefill and decode phases of LLMs by tailoring computation formats, weights, and storage placement. On Qwen 3 and Gemma 3, removing activation quantization during decoding improves accuracy on decode‑heavy tasks without increasing inference cost, while training separate compute‑native prefill weights accelerates prompt processing compared with weight‑only inference. The approach achieves faster, more accurate prompting and generation by decoupling the two phases.
RayOrch is a programmable system for lineage‑controlled multi‑grain dataflows for foundation‑model data preparation. It enables scalable pipelines that transform heterogeneous documents and videos into structured records while expanding each parent item into an ordered, input‑dependent sequence of children; GPUs batch across parents without losing lineage, child order, completion status, or routing. By addressing the limitations of coarse‑grained jobs or flat records, RayOrch provides a dataflow platform that preserves lineage and supports complex scheduling for data‑intensive model training.
PISA introduces block‑sparse attention with pyramid Top‑K selection to achieve log‑linear complexity for long contexts. By gradually pruning candidate key blocks across multiple levels, it avoids scoring all query–block pairs and reduces the attention cost while preserving accuracy. The result is scalable long‑context attention suitable for large language models with improved efficiency.
InternW0-Δ is a unified World Action Model pretrained on a heterogeneous open-data corpus to jointly model predictive dynamics and action generation for generalist robot manipulation. It integrates pretrained visual dynamics, scene semantics, 4D geometry and motion priors, and an action generator into a single framework and outperforms prior methods on simulation benchmarks and real-robot platforms. Trained on 20K+ hours of open data, the model demonstrates robust cross-domain generalization and coherent action execution.
HoTS introduces Homophily-Aware Temperature Scaling for calibrating GNN predictions. A global logit temperature is insufficient, so HoTS adapts temperature using local homophily to improve calibration under varying neighborhood structures.
Efficient Dynamic Algorithms for Graph Neural Networks with Non-Linear Propagation develops dynamic maintenance algorithms to update node representations under edge insertions for non-linear propagation models. The results enable faster, scalable updates on evolving graphs.
Attention Graphons provides a graph-limit perspective on graph transformers. By viewing attention matrices as finite samples from an underlying attention kernel (a graphon), the work analyzes concentration and stability of attention patterns under increasing graph size using dense graph limit theory and cut-distance tools.
Two-Sample Testing for Inhomogeneous Random Graphs in Non-Integral L_r Norms studies the problem of testing equality of edge probabilities between two populations of graphs under non-integral L_r norms r > 2, addressing a gap between known upper and lower bounds and proposing a test with competitive rates for aligned graphs.
The paper proposes a graph-theoretic approach to latent structure learning for nonlinear and dependent factors, using pairwise dependence measures to identify the number of latent factors and the nonlinear mapping between latent and observed variables. It demonstrates identifiability of both quantities under certain conditions.
TwinS-GCN introduces a spectral conjugate for spectral graph convolutions that injects directionality via a skew-symmetric operator. This helps avoid oversmoothing and enables modeling long-range dependencies.
Making LLMs Truly Forget proposes a deep unlearning framework that searches, selects, and severing knowledge paths. It addresses the vulnerability that facts can be recovered via multi-hop reasoning even after forgetting direct evidence. The method integrates exploration of explicit outputs and latent reasoning traces to achieve true forgetting.
Signal or Noise? Multimodal GraphRAG studies how different modalities contribute in retrieval-augmented generation. It questions the assumption that more modalities always help, showing redundancy can distract the model, and investigates each modality's contribution across questions.
Equivariant Neural Primal-Dual Assignment (ENPDA) learns a shared matching policy for Maximum Common Edge Subgraphs that generalizes to new graph pairs without retraining. This equivariant approach enables efficient MCES queries, particularly for molecular similarity search.
Transfer Learning for Edge Classification on Dynamic Text-Attributed Graphs formalizes a leave-one-domain-out (LODO) transfer protocol and shows that state-of-the-art self-supervised methods struggle under distribution shifts.
Edge-Level Automorphism in GNNs provides a quantitative framework to study automorphism at the edge level and designs to mitigate the node automorphism problem. It shows that standard GNNs collapse automorphic nodes and harms link prediction.
Reference-Tail Trust provides certified probability floors for learned updates inside deployed GNNs. It allows learned updates against a frozen backbone by constraining worst-case cross-entropy increases and using trajectory-validated tubes with an independent optimizer. The framework enables reliable, verifiable updates in live graph systems.
Before Answering investigates evidence sufficiency under memory constraints. It shows that memory-size leakage can reveal labels in evidence-sufficiency benchmarks and introduces a size-matched memory construction to prevent leakage, validated on MuSiQue.
Graph Memory proposes a spectral dense associative memory for graphs, storing and retrieving relational patterns via a Dirichlet-energy-based objective. Retrieval is performed through a log-sum-exp energy tied to the graph spectrum.
We introduce the Active Causal Discovery Benchmark (ACDB), an SCM-grounded environment for evaluating whether LLM agents can recover causal graph structure from observations under budgeted interventions. ACDB pairs a linear-Gaussian world generator with a fixed observe-intervene-submit API and a three-layer scoring contract for skeleton, DAG recovery, and intervention efficiency. Six-level evaluation shows PC with a greedy active orientation heuristic as the strongest non-oracle method.
The paper investigates the robustness of zero-shot graph models (ZGMs) under adversarial perturbations on unseen graphs. It benchmarks how transfer to target graphs holds when those graphs are adversarially manipulated, highlighting gaps in robustness evaluation for ZGMs. The work provides a framework to assess adversarial resilience in zero-shot graph reasoning.
Non-Adaptive Learning of Sparse Erdős–Rényi Graphs via Affine Splitting analyzes non-adaptive edge-detection queries for recovering sparse graphs. It shows that, for n-vertex graphs with k edges, worst-case query complexity is Omega(min{k^2 log n, n^2}) even with small error tolerance.
We build a decoupled analytics framework with a synthetic GNN benchmark in which label noise and feature distribution shift can be varied independently. Across 41 configurations and 410 runs, we study their separate and joint effects on robustness to guide robust GNN design.
Spectral Reversal identifies a spectral bias in self-supervised pre-trained GNNs when used with graph prompting. The method counteracts this bias by reversing the spectral directions to improve downstream prompting performance.
Learning the Graph and the Embedding Together proposes an affinity-guided rewiring method that jointly estimates the graph and node representations in an EM-like alternating procedure. The method is classifier-independent and aims to improve heterophilic node classification.
This work asks whether architecture-level conclusions about heterophily robustness hold when node representations vary. By creating parallel feature variants for two heterophilic benchmarks, Roman-Empire and Amazon-Ratings, comparing different representations, they show that robustness conclusions depend on the input representation. The findings emphasize that representation choices can alter heterophily robustness.
A Unified Uncertainty Representation for Graph Neural Networks via Doubly-Spectral Stochastic Expansion models uncertain node embeddings as random graph signals. Graph Fourier filters capture structural variation while a scalar orthogonal-polynomial chaos coordinate encodes latent stochastic variation, yielding a single representation for calibrated predictions and robust task-specific readouts.
HyperReCo advances evidence retrieval for LLM multi-hop reasoning by using hypergraph neural networks to retrieve and connect evidence. Unlike methods that rely on predefined expansions, HyperReCo preserves higher-order entity associations and enables query-dependent interactions.
We propose Extremely Fast and Compact Binary Graph Representations via Randomized Operator Sketching. The method constructs binary node representations directly from graph structure in an algebraic, feature-free manner, enabling ultra-fast hashing with little memory.
Efficient Message Passing for PDE Priors presents probabilistic inference on factor graphs to incorporate PDE priors into learning. The method uses message passing to approximate the posterior over parameters given PDE constraints and observed data.
This work studies how AI systems reason over patient knowledge graphs when physiology-driven signals collide with clinician decisions. It introduces ClosedLoopBench, a benchmark with 29 relation types whose signs are fixed by physics, pharmacology, or clinical practice, built from 3,442 VitalDB surgical cases with negative-control action streams. The study investigates robustness by replacing each patient’s actions with alternative policies to isolate the effects of feedback loops on the signals.
CalibHyper introduces chance-corrected relational hypergraphs for few-shot molecular property prediction. It models the joint label distribution and corrects for marginal dependencies by subtracting an independent baseline. This leads to improved accuracy by capturing dependencies among properties.
SafeMol identifies jailbreak vulnerabilities in molecular multimodal models under text-only and graph-conditioned inputs. It argues safety must be robust across modalities, balancing safety, refusal, and utility. To address this, SafeMolBench is proposed, a benchmark with 3702 samples covering 618 unique hazardous molecules to evaluate multi-modal safety alignment.
WorldGraph proposes graph-native world modeling, where evolving graphs themselves constitute the world model. It treats graph evolution as the dynamic environment to be modeled and predicted, enabling agents to reason about relational dynamics directly on graphs.
FAST-Brain presents a flow-aligned spatio-temporal surrogate brain model for rs-fMRI. It addresses long-horizon temporal dynamics, anatomical spatial structure, and high-dimensional ambient signals by a unified flow-aligned generative framework. The model enables fast, accurate brain activity simulation.
ConRAG formalizes the task of multi-hop relation inference across documents: given two endpoint entities, it aims to recover bridge entities and evidence-grounded reasoning chains that connect them, across a corpus, and to generate explanations grounded in text.
HyperMCTS augments Monte Carlo Tree Search with a hypergraph that captures recurring decision groups across long-horizon tasks. This enables reuse of trajectory feedback and more efficient search for LLM agents, reducing compute while improving solution quality.
ViCoR studies selective structure recognition for optical chemical structure recognition and proposes spatially aligned verification and executable revision to automatically produce reliable outputs while rejecting uncertain ones.
STITCH-RAG introduces a hypergraph-based retrieval framework for multi-hop RAG. It preserves topic-level co-participation and per-occurrence entity descriptions by tracing spatio-temporal influence over documents, addressing gaps in chunk-based retrieval and unlabeled projections.
HyperLabel is an encoder-decoder framework that explicitly models high-order label dependencies via hypergraph neural networks for multi-label classification. The approach improves prediction performance by incorporating structured label correlations.
Scalable GNN-based Knowledge Graph Representation Learning proposes scalable methods for KG representation with efficient message passing. It addresses scalability bottlenecks beyond subgraph sampling and aims to preserve expressivity while enabling training on large KGs. The approach improves scalability without sacrificing performance.
GeoF learns propagation geometry from message-passing feedback. It maintains a local geometry per node, initialized from a structure-aware atlas, and evolves node features and propagation geometry jointly. The recurrent framework enables adaptive, geometry-aware message passing.
Controllable GNN Explanations via Multi-Metric Preference Selection proposes optimizing GNN explanations across multiple metrics, including fidelity, interpretability, sparsity, and stability, to reveal trade-offs and enable user-driven explanation preferences.
HARMONIA presents interpretable graph learning via mixtures of neural bases, addressing scalability and flexibility limits of previous interpretable graph Additive Models. It provides explicit source-to-target contribution decomposition through neural bases and enables scalable, interpretable graph learning.
This paper argues for a selective neuro-symbolic reasoning approach to autonomous driving question answering, distinguishing queries that have exact symbolic solutions from those requiring interpretation. It introduces a query-adaptive framework that allocates computation between symbolic and neural components, leveraging a hierarchical structure to route each question accordingly.
TemporalGraphLLM proposes integrating Temporal Graph Neural Networks with Large Language Models to handle dynamic text-attributed graphs (DTAGs). The approach aims to jointly model temporal evolution of graph structures and textual attributes, leveraging both structural dynamics and language understanding.
Topology-Adaptive Hyperbolic Graph Attention Networks (HSO-GAT) introduce Hyperbolic Sombor Index as a structural prior to guide attention in hyperbolic space. This prior enables topology-aware attention that better captures hierarchical patterns.
Algorithmic Harms Associated with Generative Model-Augmented Recommendation Systems analyzes potential harms when generative models are integrated into recommender systems. It extends harm taxonomies to cover endogenous harms such as sanitization and misalignment and discusses mitigation strategies.
AutoHGNN is a neural architecture search framework tailored for hypergraph neural networks. It extends the search space with a Hyper-Interaction Module to better capture higher-order relations and mitigate topology–interaction mismatches. The method yields robust and efficient hypergraph architectures for various tasks.
The paper studies how k-NN based preprocessing shapes fairness in graph neural networks. Graphs constructed by kNN are evaluated with equalized odds as a fairness criterion, and a fairness-driven loss is added to optimize. The results show how preprocessing choices affect equalized treatment across groups.
DynGraphAgentBench is a benchmark for agentic lifecycle control in dynamic graph anomaly detection. It provides seven temporal graph datasets, eleven detectors, and eight deployment windows, with a controller granted only time-causal aggregate context in each window. The benchmark enables executable evaluation of decision-making under delayed feedback.
T-SNN introduces Temporal Simplicial Neural Networks for EEG decoding by modeling recordings as evolving simplicial complexes. It combines simplicial convolutions with recurrent updates to capture higher-order interactions among brain regions. The approach improves decoding of brain states from EEG data.
M3OS presents a Monte Carlo Graph Search–Orchestrated multi-agent LLM system for evidence-traced molecular optimization. It decouples molecular-design reasoning from optimization-state management via persistent graphs linking evaluations. A Monte Carlo graph search guides decision making across agents.
This paper tackles spatial indistinguishability in spatiotemporal prediction, where nodes with similar history diverge in future, hindering forecasting. It proposes an optimal transport–guided masking strategy to identify and mitigate such ambiguous nodes, improving predictive accuracy in sensor networks.
Temporal Heterogeneous Graph Pretraining studies how explicit encoding of two temporal signals—record ages relative to a prediction cutoff and fixed inter-record intervals—affects relational deep learning. The authors propose a unified pretraining framework that combines multiple temporal representations to improve downstream tasks.
TreeRef-BFN introduces an equivariance-free de novo molecule generator based on 2D topology and internal 3D geometry. It uses a tree-based representation to handle variable molecule size and supports conditional tasks like fragment completion and scaffold decoration without assuming fixed-size coordinates or rigid equivariance.
GT-PSSM offers a unified probabilistic framework for modeling stochastic dynamics and dependencies in multivariate time series for anomaly detection. By treating dynamics probabilistically rather than deterministically, it improves robustness to noise and structural variability in real-world data.
EngramRAG proposes a dynamic usage-weighted topology and synaptic consolidation memory architecture for multi-hop agentic memory in autonomous LLM agents. It integrates a low-latency waking state with an asynchronous dreaming state to address associative blindness, scaffolding amnesia, and static topology stagnation.
From Anomalies to Failures constructs causal error graphs for agentic trace diagnosis in LLM-driven agents. It distinguishes anomalies, errors, and failures and provides a structured causal framework to trace how errors propagate and amplify into task failures.
The Adaptive Consistency Graph (ACG) is proposed to preserve coherence during long-horizon agent execution. It incrementally organizes execution evidence and provenance to prevent later decisions from deviating from the original objective, improving alignment over long sequences.
A training-free symbolic-probabilistic framework named CKG Reasoner maps patient observations to explicit medical knowledge graphs. It integrates evidence features, patient-reference matching, a bounded information gate, knowledge-weighted evidence accumulation, disease similarity, and decisive rules, with mechanisms to handle missing evidence and auditing coverage.
We introduce EEGAgentBench, a benchmark to evaluate LLM-driven EEG analysis agents across short- and long-horizon tasks. The framework emphasizes iterative evidence accumulation, multi-step reasoning, and tool use, and provides standardized protocols to assess agents' reasoning over time.
The paper proposes reliability-informed correction with event graphs (RICE-Alpha) for LLM-agent stock forecasting. It encodes issuer-specific chronology and information availability via event graphs to guide predictions under point-in-time constraints and detect when historical transitions contribute information beyond the current forecast.
MegaGraph tackles efficient training of large-scale Graph Transformers using automated hybrid parallelism. It targets memory bottlenecks from attention score and topology-aware bias and workload imbalances from embedding layers. The approach combines system and algorithmic strategies to scale Graph Transformers on large graphs.
E3J is a fast Euclidean equivariant backend for geometric deep learning with JAX bindings on GPU and TPU. It uses optimized CUDA and Pallas kernels along with algorithmic improvements to achieve state-of-the-art throughput in forward and backward paths, including a substantial speed-up over cuEquivariance on MLIP tasks, while remaining fully open source.
CollisionGAT provides controller-agnostic one-step collision screening for multi-agent motion. A graph-attention network reads current and proposed states of moving agents along with nearby obstacles to output a collision-risk score per agent, usable by any controller to accept, repair, replan, or delay steps.
Graph-Based Learning for Multi-Horizon Martian Atmospheric Forecasting introduces MaGMA, a graph-based data engineering framework that converts OpenMARS reanalysis fields into structured learning objects. Local atmospheric patches are represented as graph nodes linked to capture spatial and temporal dependencies for multi-horizon forecasting.
SATURN is introduced to model crystallized intelligence from 21-day wearable actigraphy data. The model uses daily summaries from Fitbit records to capture sleep–activity patterns and environmental interactions via a temporal graph regression network, enabling prediction of adolescent Gc.
We propose an attention-driven heterogeneous GNN for credit card fraud detection. The model captures evolving fraud patterns in imbalanced data by modeling heterogeneous transaction relations and using attention to highlight informative signals.