Graph Learning · LLM × Graph · Multi-Agent · Science
Showing 26 papers for 2026-09-02
Nothing cleared the bar today. Papers read in full today: 6.
The load-bearing idea is not a new MLIP but the multi-layer completeness argument: under genericity, overlap, and connectivity assumptions, local 3-body messages can propagate sufficient information to reconstruct a complete representation of an L-hop environment. The paper does deliver a substantial formal chain—completeness iff UAT, switched-Gram constructions across cutoff strata, DeepSets/MLP realizability, and architecture corollaries—but supplies no empirical tests of whether its assumptions hold for real materials or whether the predicted distinction (MLP-based CHGNet/DPA3 versus linear-gated ALIGNN) matters in practice.
Its guarantee excludes non-generic equal-radius same-species neighborhoods and additionally needs overlap/connectivity conditions, so it does not directly cover many symmetric crystals or establish practical accuracy, and the DPA3/CHGNet simulation details are deferred to unavailable supplementary material.
The load-bearing result is a useful negative diagnostic: none of the tested fair-LP methods is robust across the broader structural-bias space, and statistical-parity disparities remain associated with power-law-exponent ratio, neighborhood heterogeneity, and information unfairness after controlling for assortativity. The paper does test this claim more directly than ordinary benchmark papers do, using thousands of generated graphs, partial rank correlations, and a block-permutation test; its taxonomy and generator make it a solid targeted read for fair graph learning, rather than a routine new-model paper.
The evidence is still largely generator-dependent and associational: varying only group imbalance and homophily does not independently intervene on the reported beyond-homophily measures, while the random-forest predictability and partial correlations may reflect their shared dependence on the same generative parameters; moreover, the classical LP baselines are mostly Node2Vec/SVD/GCN-era rather than a broad modern LP suite.
The load-bearing idea is unusually clean: shortest-cycle length and multiplicity provide a compact alternative to a bounded cycle dictionary, and the paper holds architecture and bond features fixed to show that on ZINC a 3-channel edge-girth feature reaches 0.0932 MAE versus 0.1005 for cycle counts through length 8 and about 0.204 with no structural feature. More importantly, it proves a concrete failure condition—edge-girth-regular graphs reduce any such edge-feature MPNN to the 1-WL bound—and verifies it exactly on all 90 applicable BREC pairs; it also candidly finds that EGAGNN is virtually no better than directly hashing the edge-girth multiset on BREC, so the architecture is not the contribution.
The practical claim rests on one supervised dataset and one target, while the best cutoff-8 cycle dictionary nearly matches the feature and several headline Table 1 comparisons are confounded by EGAGNN's access to bond attributes that GatedGCN/GIN/GCN lack.
The load-bearing idea is diagnostic rather than architectural: candidate recall/composition, pool size, off-candidate scoring, and decoding randomness are confounded with claims about an LLM reranker. The paper substantively tests that claim with matched-pool runs, retrieval recall and oracle ceilings, strict off-pool zero credit, and repeated temperature samples; its concrete finding is that proprietary LLMs beat EASE on the semantic c250 pool, but EASE/SASRec candidate pools lift both tested LLMs by 52–59%, while all tested open-weight models remain below EASE under that protocol.
The evidence is limited to one small, old movie-dialogue benchmark, and the paper itself admits that its model-family comparisons do not equalize input information, while its retriever effects do not disentangle target availability from candidate composition and prompt compatibility.
EDGE presents an Error Dependency Graph-guided multi-error attribution framework for multi-agent LLM systems, building an error dependency graph from observed failures and validating a reliable causal subset via counterfactual rollout, followed by a two-stage LLM-as-judge detector.
ISO-RAG introduces Isoperimetric Noise Control for Retrieval-Augmented Generation, addressing semantic drift and latency in multi-hop reasoning. It projects knowledge graphs into a hyperbolic Poincaré ball to guide efficient precomputation and retrieval.
MUGEN generates unlearnable graph examples to protect learning across multiple tasks. Unlike prior methods that target a single downstream task, MUGEN perturbs graph data to degrade generalization for diverse objectives such as node classification, graph classification, and link prediction.
H2Table introduces hierarchical hypergraph-enhanced large language models for complex table reasoning by representing tables as hierarchical nested hypergraphs and employing a tailored hypergraph encoder.
AdaptNTK proposes adaptive uncertainty quantification and active learning for neural network potentials to balance computational cost and reliability. It expands the training set by selecting uncertain configurations while addressing redundancy in acquisition batches.
We provide an expository introduction on the importance of higher-arity tensor operations to deep learning, followed by a novel empirical investigation of higher-arity phenomena in trained networks. The work also introduces a hypergraphical generalization of the multilayer perceptron and explores connections to evolutionary algorithms, ending with promising directions for future research.
CATeye introduces Coupled Attribute-Topology Invariance Learning to tackle voucher abuse detection under distribution shifts. It addresses coupled attribute-topology shift where edges built from attribute proximity induce environment-driven attribute changes, and learns invariant representations robust to regional and temporal shifts.
ValueGraph is a graph pre-training framework that uses automatically inferred moral-value signals as noisy auxiliary supervision to contextualize user representations. It aims to capture value-oriented tendencies in online discourse to enrich contextualized user embeddings.
EGT-KG proposes Evidence-Grounded Typed Knowledge Graph retrieval to improve scientific QA with small language models. It leverages typed knowledge graphs as structured evidence to enhance information retrieval under limited literature and context.
DREAMS investigates structured context modeling for conversational recommender systems using dual-node Monte Carlo Tree Search, introducing elicitation and exploitation node types and guiding elicitation with MCTS to better track user preferences over turns.
Athena introduces a graph-based approach to identify affected libraries for known vulnerabilities. By modeling vulnerability databases as a knowledge graph and applying graph completion, Athena can infer missing or incorrect library-vulnerability links, addressing gaps left by traditional text-based methods.
This work tackles academic question answering by reasoning over heterogeneous scholarly graphs (author, paper, venue). It introduces Agent-Enhanced Heterogeneous Graph RAG, which adapts retrieval strategies to query complexity, adds sufficiency evaluation to avoid incomplete evidence, and provides structured verification against graph facts for more reliable answers.
We present an additive, second-pass approach to uncover latent relationships in a document-derived property graph. Each document is chunked and embedded once; top-k nearest-neighbor queries across chunks yield candidate node pairs, scored with Shepard inverse-distance weighting on a rescaled chord distance metric.
EM^2Mem introduces an event-centric multimodal memory for LLMs that binds heterogeneous evidence to event anchors during memory construction, addressing the fragmentation and attribution challenges in long-video question answering.
Automated Tree Knowledge Graph Construction uses ontology expansion and retrieval from Vietnamese history textbooks to build end-to-end KG pipelines for a low-resource language, and evaluates hierarchical retrieval strategies.
Verifiable Disaster Storylines and Causal Knowledge Graphs present a citation-grounded pipeline that fuses EM-DAT disaster records with ReliefWeb and EMM documents to produce source-grounded disaster storylines and causal knowledge graphs for responders and analysts.
Delegation Without Trust argues for evaluating agent security under an untrusted-model assumption, highlighting gaps in identity, authorization, and runtime governance, and advocating robust, default-trust designs to guard against prompt-injection and other threats in multi-agent LLM systems.
The zbMATH Open Knowledge Graph presents a large-scale RDF knowledge graph spanning over 250 years of mathematical research, integrating expert-curated semantic content such as reviews, keywords, classifications, software references, and disambiguated authorship to enable advanced analyses.
Tencent Generative Recommendation (TGR) advances industrial recommendations toward a generative paradigm. It unifies generation and reasoning with components like GenRank and CCFormer to enable unified feature tokenization, scalable transformers, and knowledge-driven ranking.
NeuroGraph is a graph-driven neuro-symbolic framework for explainable threat reasoning in advanced manufacturing. It combines graph-based retrieval augmented generation with ontology-consistent multi-hop reasoning to provide transparent evidence tracing and reduce hallucinations in cyber threat intelligence workflows.
Towards Agentic Cloud Engineering introduces an agentic AI framework that transforms natural-language cloud-engineering tasks into constrained, verifiable workflows. It enables graph and loop engineering with a zero-trust agent harness to manage workflow progression, execution, and verification.
We present DUD, a memory-efficient algorithm to discover unique column combinations on disk-resident data. By exploiting the link between UCC discovery and transversal hypergraphs, DUD avoids quadratic difference-set generation and uses targeted pruning to scale to large datasets with limited memory.