Graph Learning · LLM × Graph · Multi-Agent · Science
Showing 25 papers for 2026-10-02
Nothing cleared the bar today. Papers read in full today: 10.
The load-bearing idea is real: the paper shows that a stronger model cannot fix missing distinctions, and that a harness must own the state that determines admissibility; that part is formalized with an exact aliasing theorem and then echoed by several targeted failures (budget overspend, stale reads, hidden ontology collisions). The evidence is mostly diagnostic rather than broad-scale, but it does test the claim it makes: the failures disappear when the missing state is owned at commit or reintroduced into context, and current-state-only checks are shown to miss stale or loop-level inconsistencies.
The empirical scope is narrow and partly synthetic, so the paper supports a framework more than it establishes a general empirical law for agent systems.
The central idea is load-bearing: for state costs, solve the generalized bridge exactly by tilting the generator, then run plain Schrödinger iterations against that tilted reference. The experiments do test that claim rather than merely showing benchmark wins: on assignment it recovers optimum assignments in the entropy band where the learned method was evaluated, on chignolin it lowers the barrier monotonically, and on the road network it matches endpoints and scales to millions of nodes with linear memory; it also includes a negative/diagnostic result that the learned TD-based approach can misplace a pair at the largest assignment size and that the regularized objective itself can fail at high entropy.
The main limitation is that the exact reduction only covers state-only running costs with an explicitly known finite reference generator; congestion and other marginal-dependent costs still need an outer fixed-point iteration, and the broader contribution is a clean synthesis of classical ingredients rather than a fundamentally new primitive.
This work analyzes On-policy distillation (OPD) for language models and how extrapolating RL-induced representation residuals can benefit a student. It shows that most of the beneficial change encoded in the teacher’s hidden states is attenuated when mapped to the LM head, and that relying on sampled-token log-probability ratios introduces noise for output-space extrapolation. Consequently, the teacher functions as a direction rather than a destination.
UniEvo-VL presents a self-evolving, on-policy training recipe for multimodal models that learn from their own constructive self-critiques during test-time compute. Rather than relying on a larger external teacher, the model acts as both teacher and student with different contexts, and the student only sees the vanilla question while benefiting from self-critique feedback.
This paper identifies co-cheating as a failure mode in self-evolving search agents, where proposer and solver converge on shared errors so internal rewards improve without corresponding external correctness. A post-hoc audit against source evidence shows co-cheating grows over successive rounds and pseudo-label accuracy stagnates or declines, despite stronger in-loop training signals. The authors discuss direct mitigation strategies to break the feedback loop and align internal signals with external correctness.
This work reevaluates latent visual reasoning (LVR) by grounding latent reasoning in visual evidence. It analyzes latent-token behavior and finds a latent evidence-credit gap: latent tokens respond weakly to image perturbations that alter the correct answer. The authors argue this stems from insufficient explicit supervision linking latent steps to visible evidence and propose grounding techniques to address it.
AREX-2 advances self-improving agents by enabling long-horizon iterative refinement at test time. It relies on two complementary capabilities—reflection, which produces a better solution, and long-horizon execution, which sustains effective iteration across many rounds. The framework is designed to be domain-agnostic and shows how to synthesize long-horizon improvement trajectories from machine learning and algorithmic sources.
WOMBAT introduces a whitebox benchmarking suite for GNN explainers by constructing 14 whitebox GNNs whose message-passing weights are hand-tuned to respond to specific SMARTS motifs. Because each model's decision rule is known, it provides ground-truth attributions to diagnose when explanations are wrong due to explainer failure vs model shortcuts. Evaluations on millions of PubChem molecules demonstrate its utility.
CrossGMN introduces Graph Metanetworks to enable weight-space transformations across architectures, addressing weight-space symmetries that complicate cross-architecture transfer. The metanetwork learns transformations that map a source model to a target architecture, enabling tasks like model editing and cross-architecture property prediction. This broadens the scope of weight-space equivariance beyond architecture-preserving transforms.
DeFA proposes a dependency-guided failure attribution framework for LLM agents, constructing an event dependency graph that fuses protocol relations and semantic dependencies. It identifies events violating task requirements and traces their sources and propagations to build a failure propagation graph for diagnosis.
We study graph neural combinatorial optimization under a hard information horizon, where each node commits to its share of a global solution seeing only its k-hop neighborhood, and those commitments must compose into a globally feasible solution. We formalize this as local set cover and instantiate on a weighted multipoint relay (MPR) selection problem arising in OLSR v2, showing feasibility and performance trade-offs.
We propose higher-order positional encodings for graph representations by lifting graphs to simplicial complexes to capture higher-order interactions beyond edges. This complements existing graph-based encodings and integrates topological information into the representation, albeit with computational considerations that we discuss.
Memetic Trojans define a new class of network-mediated adversarial payloads in autonomous LLM agent ecosystems. They exploit agents' tendency to retransmit and amplify content by embedding malicious payloads in memetic content, enabling endogenous propagation rather than externally induced infections like agent worms. The paper analyzes the threat model and motivates defenses against contagion-driven attacks.
We argue that Semantic IDs are best viewed as recursive clustering over a graph with semantic content and collaborative edges; GrIS reframes SID construction as hierarchical graph partitioning, uniting semantic and collaborative signals under a graph-informed framework for improved generative recommendations.
We study whether tissue-specific interactomes carry a reliability signal for GNN predictions. Using effective resistance as a proxy for information flow, we find a strong inverse-degree bias across 24 tissues and that reliability degrades as the co-expression filtered network grows, with a Spearman correlation of -0.955. The residual deviations from this limit exceed degree-preserving null baselines, highlighting the need to account for topology when interpreting predictions.
We isolate the transport primitive in Sheaf Neural Networks via quiver representations and connect it to multi-head attention. This shows that standard multi-head attention implements diagonal edge maps when viewed as local transport coordinates. We discuss extending beyond diagonal maps to enable cross-head interactions for richer message passing.
EviGraph provides proof-carrying selective recommendations over temporal public-service knowledge graphs. A language agent maps decision requirements to evidence in a temporal knowledge graph, while a deterministic checker ensures recommendations are supported by the evidence. Evaluations on bilingual HK benchmarks show effective, explainable recommendations.
We study dynamic reasoning with revision-aware independent agents, where events revise task bindings and the system must select the valid document version for each query. We adapt six benchmarks into 31k dynamic episodes to analyze how agents handle revisions and maintain consistency.
MatRAG introduces a Matryoshka Hierarchical Retrieval-Augmented Generation framework for efficient multi-hop QA. By combining RAG with Matryoshka Representation Learning, it reduces indexing and query costs while preserving retrieval quality through hierarchical representations.
MOVE introduces a framework for open-world verification and expansion in graph learning. It tackles post-deployment emergence of new classes and goes beyond unknown-node detection by evaluating whether the existing label space is sufficient and whether generated class descriptions are reliable for expanding the taxonomy. The authors show that multimodal signals beyond individual modalities are needed for accurately inferring unknown nodes and outline a verification-expansion pipeline.
This work analyzes federated multimodal GFM adaptation and argues for coupling perception and reasoning by jointly updating both the multimodal encoder and the graph-side GNN across clients, instead of freezing the encoder. Such joint optimization addresses privacy constraints in decentralized graphs and improves adaptation to heterogeneous, open-world data.
We propose Higher-order Grammar Representation (HGR) to encode molecules as combinatorial complexes, enabling explicit representation of rings and motifs for generative and foundation models in chemistry. HGR provides topology-aware representations that improve expressivity while remaining tractable for generation and downstream modeling.
BranchIP introduces Branch Interatomic Potential, a single-model framework for adaptive, learned tensor-product computation in equivariant MLIPs. It uses a novel distillation loss to learn when to apply expensive tensor products, achieving accuracy while reducing computation for interatomic potentials.
We present Varda-single-1.0, a deterministic, data-driven weather forecasting system for Switzerland at 1 km resolution. It uses two graph-transformer based models in an encoder-processor-decoder setup: a 6-hourly autoregressive forecaster and a temporal downscaler to reconstruct hourly fields, trained in a staged curriculum.
The paper proposes a graph-based method to detect coordinated disinformation campaigns on Telegram. By starting from verified disinformation sources and modeling account interactions as a graph, it uncovers previously unknown accounts that participate in spreading false information, revealing hidden coordination.
To enable scalable patent prior art search, the authors move beyond brittle rule-based parsers and introduce a neural parser that predicts invention graphs directly from patent text. They adapt biaffine attention from dependency parsing with a locally constrained scoring mechanism to build structured patent graphs for retrieval.
We propose a framework to learn ab initio phase-field models by projecting molecular dynamics onto density fields via Mori-Zwanzig, deriving nonlocal evolution laws for mesoscale microstructure. This approach aims to achieve quantum-mechanical accuracy at mesoscopic scales by learning effective free energies and mobilities.
Build2SPARQL introduces a large-scale NL-to-SPARQL benchmark for building knowledge graphs, enabling natural-language querying over Brick and ASHRAE 223P ontologies. It fills a data gap by providing a scalable dataset to train and evaluate text-to-SPARQL systems.
BuildGraph is a synthetic building knowledge graph dataset with 120 buildings in Brick Schema. Grounded on DOE prototype models and real building sensor-placement patterns, it covers eight commercial building types across three ASHRAE energy-code vintages, enabling public semantic querying.
We present ECMO, an OWL-based modular ontology for cross-sector crisis management, with ECMO-CORE capturing core concepts like hazard, event, exposure, impact, and response. The ontology uses punning to resolve ambiguities between hazard types and event manifestations and includes modules for domain-specific crisis management.