Graph Learning · LLM × Graph · Multi-Agent · Science
Showing 44 papers for 2026-09-09
Nothing cleared the bar today. Papers read in full today: 10.
The load-bearing idea is not a new architecture name but a controlled diagnostic: embed a known solver exactly in the initialization/edge-message state, then test whether learning improves the solver rather than a weak neural baseline. The full text does support the narrow claim: min-sum-GNN substantially improves non-decimated and decimated min-sum on random-regular, Watts-Strogatz, and several Ising ensembles, whereas ordinary and spectral GNNs often fail; importantly, O(n log n) simulated annealing reliably remains better. This is a useful negative/result-oriented paper for learned optimization, though it is an incremental specialization of neuralized BP rather than a general breakthrough.
The headline practical conclusion is based on a narrow collection of synthetic Ising graph distributions and mainly compares against simulated annealing rather than strong dedicated MaxCut/Ising solvers (for example SDP-based or modern specialized local-search methods), while the best-of-10 trial protocol and CPU-versus-GPU timing further complicate cost comparisons.
The load-bearing idea is the distinction between flat procedural memory and local, connected transition context: the same graph performs substantially better when guidance is generated from a localized subgraph than when the full graph is injected or summarized, especially on GDPval and ALFWorld. The paper also supplies a meaningful diagnostic result: a hand-authored graph can severely hurt MultiChallenge, while iterative validation-gated evolution recovers it, and the cross-model table is broader than a single benchmark win. However, the evidence establishes that this particular prompting-and-workflow assembly helps, not that graph topology itself—rather than better localized textual instructions and an additional LLM call—is the essential cause.
The adaptive evolution claim is fragile: accept/reject decisions on EnterpriseArena use only 20 validation episodes, repeatedly query that validation set across rounds, and lack comparisons to equally budgeted text-only rule/prompt evolution, hard workflow/state-machine control, or non-graph local retrieval.
Diffusion-augmented LLMs define an autoregressive distribution while using diffusion to sample multiple tokens in parallel. The model decouples parameters into autoregressive weights (trained with standard next-token prediction) and lightweight diffusion weights (trained to generate several tokens at once). This design aims to achieve lossless speedups by parallel token generation without changing the underlying autoregressive distribution.
FlowBalance proposes verifier-grounded self-improvement for on-policy reasoning. It addresses the fragility of the inner loop where sparse terminal verifiers provide reliable supervision while dense same-model guidance can reinforce overconfidence. The approach learns a normalized distribution over complete responses and, for each on-policy trajectory, uses a frozen policy view with privileged context to compute token-level log-probability gains, which are aggregated into a trajectory-level objective to guide improvement.
ENEAS introduces an embedding-guided neural ensemble for adaptive segmentation and instance tracking. It targets weaknesses of text-promptable segmentation models (e.g., temporal hallucinations, spatial fragmentation, semantic misclassification) by combining multiple embeddings to improve consistency and object-level reasoning, including reporting when a target leaves the field of view and avoiding fragmentary or misled segmentations.
Causal Foundation Models (CFMs) adapt the foundation-model paradigm to causal inference. By pretraining networks to estimate causal quantities across tasks and modalities, CFMs aim to perform causal reasoning on new problems without bespoke pipelines or fine-tuning. They enable prompting-based conditioning of causal estimates, broadening applicability across tasks without task-specific retraining.
Verify Before You Distill introduces Teacher-Gated On-Policy Distillation (TG-OPD), a prompt-level gating mechanism that decides when to apply teacher supervision during on-policy distillation. It addresses the risk that a confidently wrong teacher can mislead updates under reverse KL and that distributional proxies may be unreliable. By verifying outcomes before applying distillation, TG-OPD improves reliability and sample efficiency of on-policy learning.
The paper shows that oversmoothing is not the only failure mode in GNNs; due to community structure, representations collapse quickly inside communities while inter-community separation persists, creating an echo chamber. It proposes diagnostics to detect this phenomenon beyond global metrics and discusses remedies.
We study budget-aware training-set selection for machine-learned interatomic potentials, showing that whether to prioritize structural diversity or model disagreement depends on how much data is kept. A budget-resolved comparison between coverage-based and disagreement-targeted selectors demonstrates when each approach is advantageous.
CUNO proposes curriculum and preference optimization to stabilize graph unlearning under mass deletion. By accounting for the varying importance of deleted samples and graph structure, it avoids catastrophic utility loss during unlearning.
LoGIC proposes budgeted context construction for node-level graph in-context learning with tabular foundation models. It analyzes which labeled and unlabeled nodes should form the prompt to balance utility and scalability, reducing quadratic attention while maintaining performance.
MLIP Detective is an active failure-mode discovery framework for machine-learning interatomic potentials. Starting from benchmarks, it generates physics-informed failure hypotheses and screens them with inexpensive checks to uncover failure modes beyond standard benchmarks.
TTGBench benchmarks topological evolution and semantic drift in text-attributed temporal graphs, addressing limitations of existing benchmarks that emphasize structural evolution. It provides datasets and metrics to evaluate semantic drift in dynamic graphs.
This work reframes one-shot federated graph learning by dropping the assumption that local GNN training is necessary. Under extreme non-IID settings, local training can hurt cross-client alignment, so the authors propose training-free statistical estimation to aggregate knowledge across clients.
PAGR (Proof-Carrying Algebraic-Geometric Retrieval) separates knowledge certification, latent representations, and multi-hop admissibility using quiver-, provenance-, and sheaf-theoretic tools to ground LLM retrieval.
Chimaera blends mixture-of-experts with graph foundation models, accommodating cross-task and cross-dataset graph learning. It uses graph prompts and linear GNNs as experts, with LLMs providing embeddings, and supports flexible combination strategies and extensions to link- and graph-level tasks.
We introduce a two-stage generative framework that uses a fixed-dimensional molecule-level latent representation to generate variable-size 3D molecules. The second-stage flow-matching model samples this latent vector, and an autoregressive Transformer decoder then determines the final molecular structure and size.
We present a graph-engineering approach for inference-time multi-agent LLM workflows. Instead of a fixed topology, we synthesize a task-conditioned temporal workflow graph and introduce ReActNet, which compiles a query and role-specific agents into a sequence of directed communication graphs in a training-free manner.
Skynet develops workflow-level anomaly detection for agentic AI using semantic and structural modeling to identify failures that propagate through long-horizon task graphs and tool usage.
GraphNOSE is an open-source graph transformer framework that predicts multi-label odor descriptors from SMILES for single molecules and binary mixtures. It integrates positional encodings to capture molecular locality and context, improving generalization to novel chemical scaffolds and complex odor mixtures.
We conduct a comprehensive study of candidate generation methods for vacation rental alternatives, comparing collaborative filtering, shallow embeddings, and graph neural networks under heterogeneous inventories and geographic constraints. The results show how ensemble and multi-source strategies improve relevance and diversity of recommendations.
SPIRE leverages a structural-entropy descriptor of the degree distribution to weigh client contributions in one-shot federated graph learning, enabling topology-aware diffusion generation. This topology-informed approach improves differentiation between clients and boosts performance under heterogeneous graphs.
This work proposes Bridge-Router-Adapter architecture for unified multimodal graph foundation models to avoid cross-scope context entanglement. It preserves scope-specific graph contexts and enables cross-modal fusion through dedicated adapters and routing components.
DrugReason combines knowledge graph reasoning and language evidence in a dynamic multi-view framework to reason about drug-disease relationships for repurposing. It addresses multi-hop biological mechanisms and aims to improve the reliability of candidate pairs.
HOPE addresses heterophily in open-set node classification with pseudo-extrapolation and edge-aware calibration to separate unknown classes in graphs where connected nodes may belong to different classes.
Polarity-Asymmetric Structural Calibration enhances link sign prediction by acknowledging asymmetry in structure priors under sign imbalance, improving predictions for minority and locally conflicting edges.
EdgeMem proposes an agent-memory system built around a multi-anchor hypergraph to preserve original interaction turns and organize them by complementary content, temporal, and episodic cues. This avoids repeated summarization while enabling robust retrieval.
SE-GoS advances Self-Evolving Graph-of-Skills for large-scale skill libraries by distilling historical executions into a retrieval graph that generalizes to unseen tasks, mitigating retrieval bottlenecks in large-scale LLM agents.
Evidence-Grounded Retrieval (AHLERT) automates hunt lead generation from CTI reports by extracting environment-aware, actionable leads grounded in observable artifacts and techniques.
Do All Nodes Benefit Equally from Knowledge Graphs? Adaptive Node-Aware KG Fusion for Recommendation proposes AdaKG, an adaptive weighting scheme that tailors KG signals to each node rather than applying them uniformly.
MemLoc proposes a Retrieve-Localize-Generate framework for long-term conversational memory QA to address fragmented evidence across sessions and noisy retrieved content. By selecting where to search (retrieve), extracting the relevant pieces (localize), and generating answers (generate), MemLoc unifies cross-session memory handling.
This paper introduces a multi-granularity adaptive hypergraph representation learning method called Granular-ball, which partitions the graph into granular balls to form hyperedges at multiple scales. It aims to better capture high-order relationships than fixed hyperedge definitions, and shows improved representations on downstream tasks.
Trust-But-Verify introduces poisoning-resilient locally private graph learning protocols. Building on local differential privacy, it defends against data poisoning attacks while preserving privacy, through robust aggregation and validation mechanisms.
alpha-Graph introduces an attention-infused normalizing flow for tractable graph modeling, enabling expressive yet tractable modeling of graph-structured data beyond traditional GNNs and pretraining.
ONE CYLinder provides a graph-based surrogate modeling benchmark for unsteady bluff-body flows, spanning laminar to high-Reynolds-number regimes. The benchmark includes 450 high-fidelity simulations (Variational Multiscale finite element method) and standardized protocols for long-horizon autoregressive prediction.
Preserving contextual information in cultural heritage metadata through multidimensional knowledge graphs argues that standard knowledge graphs cannot capture context-dependent validity of statements in CHO metadata. It proposes multidimensional knowledge graphs to accommodate evolving or conflicting viewpoints (e.g., colonial versus post-colonial perspectives or shifting scientific consensus) beyond provenance alone, enabling richer contextualization through multiple contextual dimensions.
Temporal Heterogeneous Graph Transformer for Credit Card Fraud Detection (THGT-FD) models transactions as a token-and-relation-typed sequence, uses Time2Vec encoding, and applies a Transformer to capture intra-transaction interactions for fraud probability estimation.
GraphFAS is a distributed system for automated graph feature generation and selection in industrial transaction networks. It uses Boruta-based feature selection and a non-parametric feature generator that builds interpretable multi-hop subgraph features for financial risk control.
TD-STGT proposes a spatio-temporal graph transformer for fine-grained mobile traffic demand forecasting. It uses a population-scaled demand proxy from crowdsourced data and daytime population to forecast demand across grids, improving planning for 5G/6G deployments.
Deposon introduces an auditable, conservation-guaranteed scattering layer over LLM reasoning paths. Each node in a concept-decomposition graph binds to a Deposon state, enabling three-channel scattering (transmission, reflection, irreversible dissipation) that obeys T+R+A=1 with machine-level energy audits.
We analyze trustworthy graph-agentic RAG systems for social good, focusing on architectures, failure propagation through the evidence graph, and assurance-by-construction methods to prevent cascading errors and ensure reliable outcomes.
Agentive Algorithm Engineering for exact minimum cuts combines an inexact solver to obtain tight bounds, bound-based reductions, improved data structures, and parallel contraction. Implemented in VieCut, it outperforms previous fastest exact methods on real-world graphs.
LEBGen presents an LLM-enhanced Bayesian network framework for few-shot travel survey data generation, modeling heterogeneous traveler groups and dependencies to synthesize representative survey records from limited samples.
Graph-Based Personalized Memory for LLM Agents details a memory framework that uses graph representations to model user-specific relations, temporal context, and evidence links, enabling personalized memory, evolution, retrieval, and evaluation.
DART proposes a DAG-based reputation and incentive framework with blockchain-enabled governance for trustworthy LLM multi-agent collaboration, balancing centralized orchestration with distributed trust and auditability.
ARNAI introduces an Artifact Removal Network for robust spinal image segmentation and measurement, integrating autoencoding and inpainting within a Restore–Segment–Measure framework to handle implants and other artifacts.
OntoKG-EQ provides a provenance-grounded, competency-question-governed knowledge graph for auditable analyst querying, ensuring reproducibility, evidence links, temporal explicitness, and validity.
X-DigCheck is a domain-independent environment for building and maintaining application profiles as they co-evolve with the data they describe. Profiles fixed to a static ontology drift from the schema they were meant to capture; X-DigCheck treats profile construction as a continuous ontology-data co-evolution loop: data are lifted into RDF against the profile, checked through competency questions and SHACL, and the resulting reports drive revisions of the ontology, mappings, constraints, and the graph.
WiDiff extracts and analyzes changes from Wikidata's edit history to understand the evolution of knowledge graphs. It helps distinguish real-world updates, error corrections, and vandalism, informing downstream tasks such as question answering, entity linking, and semantic search that rely on time-aware content.