Graph Learning · LLM × Graph · Multi-Agent · Science
Showing 13 papers for 2026-09-14
Nothing cleared the bar today. Papers read in full today: 6.
The load-bearing contribution is an evaluation reframing rather than a new editor: locality for ranked KGE answers is an editor–protected-scope property, not merely a question of whether reused parameters or selected scores changed. The experiments directly test this diagnostic claim: direct entity promotion has zero parameter-support damage yet displaces a relation-scope top-10 answer in about 77% of edits, while exact null-space preservation is safe but feasible for only about 1.3–1.4% of cases; the rank-nullity/alignment analysis usefully explains that failure. This is a solid, practically consequential diagnostic result for anyone evaluating KGE edits, though not a field-wide new editing capability.
The evidence is strongest only for single low-ranked-triple edits on FB15k-237 DistMult/ComplEx, and its relation-scope protection and blocker exclusion use retrospective observed labels across splits, so the reported safety trade-offs do not yet establish deployable behavior for incomplete KGs, sequential edits, or false-fact correction.
The load-bearing idea is epistemic rather than architectural: distinguish a KG's endpoint-discovery value from whether its intermediate path actually constrains the generated mechanism. The paper does test that claim unusually directly—endpoint-only outputs have the best overall rubric score (16.68 versus 15.14/14.96 for path conditions), while full paths modestly improve automated evidence proportionality (+0.20–0.23), human reviewers prefer path-grounded outputs on 90.3% of the 55 sampled paths, and shuffling intermediate order lowers proportionality by 0.59–0.79. This is a useful diagnostic result for GraphRAG/KG-grounding researchers, although it is a carefully executed evaluation reframing rather than a generally demonstrated biomedical discovery method.
The central "evidence proportionality" outcome is partly subjective, has the weakest judge agreement, and is vulnerable to circularity because judges are told the prompting condition and path-grounded generations are inherently supplied with more evidence; the 55-path human sample does not fully resolve that issue.
The load-bearing idea is not merely adding a chirality feature: an even atom–stereogenic-unit field chooses where a signed relative rotation acts, enabling non-anchor-to-non-anchor attention to carry handedness. The paper does test that mechanism with equal-support global-versus-anchor interventions, sign-removal controls, coordinate-reflection audits, and a scarce-mirror-supervision experiment; its strongest practical finding is the substantial axial ACMP gain over the like-for-like ChiDeK retraining (77.8% versus 65.2% Rotation; 74.5% versus 66.6% Symbol), whereas central-task gains are small or saturated. The exact ECD pair-consistency guarantee is real by construction, but it comes from evaluating both signs and projecting the representation, and the authors appropriately show that this can reduce raw fully supervised Symbol accuracy.
The evidence is compelling for curated central/axial stereogenic units on a small axial benchmark, but it does not establish broad molecular-property generalization, and the method adds O(|S|N²d) attention work while relying on a canonical-role extraction whose guarantees are conditional on tie handling.
GTA introduces Graph Theory Agent and Benchmark for Algorithmic Graph Reasoning with LLMs. It presents GT Bench, a benchmark of 24 classical graph problems across 44 task-structure settings, with over 100,000 examples in four representations (natural language, code, etc.), to evaluate LLMs’ ability to reason about graph algorithms rather than just code generation.
GraphAHA introduces Graph-Based Adaptive Search with Heterogeneous Actions for Test-Time Code Generation. It addresses inefficiencies in tree-structured search by sharing states and statistics across converging trajectories and coordinating sampling, repair, and reasoning actions to improve code generation at test time.
SNAP-KG assigns new entities to semantic communities using only their raw features, with no graph access and no retraining at inference. The authors evaluate SNAP-KG on multiple benchmarks and a 2.4M-node OGB-WikiKG2 KG, where at least one graph view is homophilous and SNAP-KG performs well. The paper then investigates what happens when homophily does not hold, extending evaluation to heterophil settings and charting the boundary of SNAP-KG’s effectiveness.
Beyond Vector Similarity introduces Hierarchical Context-Resident Graph (HCRG) for enterprise code migration. Traditional RAG relies on vector similarity and often fetches isolated chunks that break inheritance. The proposed pipeline uses tree-sitter for ASTs and a hierarchical graph context to preserve topology and inheritance, improving compilation success rates.
Cognition on Graph advocates bidirectional graph-text synergy and cognitive cycles to navigate massive knowledge spaces. It critiques reactive, graph-first exploration and proposes a bidirectional loop where graph and text retrieval inform and refine each other for complex reasoning over heterogeneous KGs and corpora.
InRTL presents a unified framework for relational table learning that explicitly models both intra-table and inter-table dependencies in PK-FK connected tables. It formalizes two complementary interaction patterns—within-table interactions and across-table interactions—to capture relational structure and improve predictive performance on multi-table data.
Repair Before Reinforce proposes a context-augmented knowledge-graph reasoning framework for multi-hop question answering. By training LLMs with KG context around facts rather than only isolated triples, the approach improves multi-hop reasoning and QA performance.
A Graph-Based Approach for Mapping Kernel-Level Telemetry to MITRE ATT&CK describes a methodology that collects kernel-level events via eBPF and builds a graph-based mapping from observed commands to MITRE ATT&CK techniques, enabling automated, threat-informed defense beyond traditional threat intel reports.
Preference-Drift-Aware Subsequence Learning and Hierarchical Context Fusion tackles long-sequence generative recommendation by addressing efficiency versus accuracy trade-offs. It proposes subsequence learning that adapts to preference drift and a hierarchical context fusion mechanism to fuse information at multiple levels, reducing noise and improving recommendations with long histories.
OneLA proposes a linear-attention decoding framework to scale large-beam generation in recommender systems. To avoid memory/traffic explosion, it exploits shared prompts and short divergent suffixes across beams, allowing a single linear-attention state to serve all beams and reducing recomputation while maintaining quality.