Graph Learning · LLM × Graph · Multi-Agent · Science
Showing 45 papers for 2026-09-15
Nothing cleared the bar today. Papers read in full today: 9.
The load-bearing contribution is an exact theorem rather than an architecture result: under regularity, finite vocabulary, admissibility, ordering, and monotonicity, “Ratchet” recurrent set-GNNs are effectively equivalent to BΣ¹◇, and Proposition 17 identifies a concrete stabilization obstruction to νY.◇Y, μX.□X, and EF EG. The full text supports the formal claim with constructive translations, checkable sufficient weight conditions via contraction/monotonicity, and a differential compiler test over 9,960 pointed-graph checks; this is a meaningful diagnostic result about what finite-state convergence cannot certify, not a benchmark increment.
The practical scope is narrow and not yet empirically validated on trained recurrent GNNs: Ratchet requires finite vocabulary plus semantic regularity/admissibility, while the stated weight-checkable tests are sufficient rather than a decision procedure for whether an arbitrary convergent model belongs to the class.
Benchmark Radar is a living database and search engine for AI benchmarks and evaluations. It aggregates LLM evaluation benchmarks, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific tests, and provides daily updates of papers, repos, datasets, and releases. It also offers a searchable benchmark catalog and captures mentions in model cards to help researchers understand the settings behind reported scores.
Feyospace-v1 proposes a data-centric framework to train frontier cyber models, addressing bottlenecks beyond simply scaling the model. It introduces five complementary systems: Choulea analyzes hidden reasoning signatures, SkyReal reduces teacher-sampling cost, Hongzwang bypasses API restrictions on teacher execution, PSBreakup restores capabilities weakened by model merging, and Kreator converts expert interventions into scalable supervision signals. Together, these components aim to lower training costs and improve the capabilities of cyber agents.
DataFlex-RL introduces an evaluation platform for RLVR data policies, enabling fair comparisons of rollout choices under a common GRPO recipe. In their experiments, 13 configurations across 12 seeds were evaluated using Qwen2.5-7B-Base across 12 math, logic, and science benchmarks, and uniform GRPO improved domain-balanced average accuracy by 7.76 percentage points over an untrained checkpoint. The study also analyzes eight rollout-selection strategies, highlighting trade-offs between policy quality and data efficiency.
Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models. We identify vision-action shortcuts that arise when models rely on spurious visual cues correlated with actions in training, hurting generalization under distribution shifts. To address this, we constrain how visual information is used to generate actions while preserving task-relevant spatial information, through Latent Interface Training to disentangle visual representations from action generation.
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking. Post-training attention sparsification reduces quadratic attention costs by selecting a small set of context units per query. Existing methods use a lightweight selector with Top-K, blocking gradients and often distilling dense attention distributions that are not directly aligned with the original LM loss. SAS proposes end-to-end optimization of context ranking to align the selector with the language modeling objective, achieving sparse attention with preserved performance.
ProtoGuide provides a post-hoc, backbone-agnostic guidance framework for class-conditional graph generation. It recovers an analogous conditioning mechanism to classifier guidance in discrete graph diffusion without embedding class information in the denoiser.
We introduce self-evolving memory for generative recommendation, a memory mechanism that evolves with user interactions to adapt recommendations without retraining, improving long-term adaptation to user preference drift.
We propose a Bayesian last-layer (BLL) extension on top of a deterministic GNN encoder to perform inductive node classification with calibrated uncertainty under distribution shift. The approach addresses the breakdown of Gaussian conjugacy caused by softmax classification and aims to provide reliable uncertainty estimates for safety-critical applications.
Cost characterization analyzes vertically partitioned federated knowledge graphs, formalizing partitioning as a design space and comparing strategies in terms of communication, indexing, load balancing, and query latency.
LiftGCN presents an energy-preserving graph learning method based on Joukowski spectral lifting to maintain high-frequency components for finite element stress prediction. This reduces smoothing and improves accuracy on irregular FE meshes.
PE-Based Deformable Graph Neural Networks propose deformable GNNs that use position-encoding-driven perturbations to adapt receptive fields, addressing depth-induced oversmoothing, long-range dependency compression, fixed neighborhoods, and noise on heterophilous graphs.
We provide finite-time analysis of persistent Gaussian perturbations in recurrent GNNs. While the global energy bound prevents asymptotic oversmoothing, it does not guarantee finite-depth node separability; the paper derives a second-moment decomposition to characterize when node representations remain distinct.
We propose a three-layer decoupled diagnostic framework to attribute errors in cloud-native Graph-RAG pipelines to data integrity issues, KG defects, or Cypher generation errors. Experiments on a spatio-temporal ecological KG show data integrity is the dominant bottleneck, with structural defects also impacting performance.
THESEUS reframes multi-hop KGQA as a question-conditioned graph navigation problem, enabling explicit tracing of intermediate reasoning steps and improving interpretability and faithfulness of answers.
We offer an evolutionary computation framework for multi-agent Q-learning with mean-field environmental feedback. Agents update stateless Q-values on a fixed graph while the population average shapes an environmental variable that modulates the payoff matrix; a deterministic transport equation for the population distribution is derived under a first-order mean-field closure.
We introduce NLKGQ, a natural-language knowledge graph query execution framework that uses controlled semantics (an OWL ontology) in the LLM context window to better transfer data-model concepts and execute queries on knowledge graphs.
HGTO offers a unified graph-based, physics-informed formulation for density-based structural topology optimization. It leverages graph representations of the finite element mesh to model density and displacement while enforcing physics constraints, enabling neural topology optimization.
GSLAD introduces prototype-regularized graph structure learning for multivariate time series anomaly detection. The method uses a two-phase training procedure to learn the graph structure and regularize it with prototypes, enabling anomaly scoring from structural deviations.
We propose a Geometric Flow enhanced graph coarsening method for GCNNs. By leveraging geometric flow, it captures higher-order mutual connections among neighbors, improving the quality of cluster-based pooling beyond purely topological information.
We present SPARC, a symmetry- and property-aware reinforcement learning framework for inverse materials design. It optimizes physical objectives while enforcing crystallographic symmetry constraints to ensure meaningful and symmetry-protected designs.
We propose GEAR, a framework that pairs dynamic social encoding with dynamic activation to control how encoded social context influences future trajectory generation, improving prediction under changing social conditions.
We develop distributed fast fixed-point algorithms for composite monotone inclusions over networks, solving 0 ∈ sum_i (G_i x + T_i x) with privacy-preserving, accelerated convergence in a decentralized setting.
This work reframes influence maximization as finding a minimum dominating set and uses graph neural networks to learn scalable, unsupervised solutions for large social networks. By exploiting graph structure, it seeks a minimal influencer set whose coverage approximates the network, addressing the NP-hardness of MDS at scale.
The study analyzes pre-training strategies for graph transformers in biochemistry. It finds supervised pre-training using computed properties as labels yields the largest downstream gains, and that constraining model capacity helps mitigate overfitting in graph transformers.
GTFD is a graph-transformer based fraud detector that fuses structural and temporal signals from a payment graph. It uses a multi-head graph attention network for structural encoding, a gated transformer for transaction sequences, and a conformal risk control module to provide calibrated risk estimates.
BSCA introduces a temporal knowledge graph embedding in a biquaternionic space, combining circular and hyperbolic rotations. A complex-valued attention mechanism adaptively fuses time-conditioned and relation-conditioned information for improved link prediction.
HiFi-Mol is a multi-view molecular representation learning framework that pretrains separate encoders for hierarchical graphs and contextualized fingerprints before fusing them for downstream property prediction.
FedV-KGQA in Practice reports empirical findings and an interactive prototype for federated vertical KGQA. Each silo trains local embeddings on its own triples, while a central server concatenates silo views to support multi-hop reasoning across partitioned graphs.
We present BusMA, a bus-style communication substrate for multi-agent systems. It enables direct peer-to-peer communication and reduces misrouting, preserving agent autonomy while supporting scalable, robust information exchange.
We introduce STHMoE, a hypergraph-enhanced method for coordinating heterogeneous dependencies in LLM-based urban traffic forecasting. It integrates temporal, spectral, pairwise spatial and higher-order cues to adapt to evolving traffic regimes.
We propose EEG-Xplain, a unified attribution framework that combines gradient-, perturbation-, and activation-based explanations to interpret EEG foundation models across spatial, temporal, and frequency dimensions. It identifies critical channels and visualizes explanations to enhance clinical trust.
We introduce PEARL, a retrieval-augmented generation based support agent for gameplay in Parallel, a parallel programming puzzle game. PEARL fuses semantic knowledge retrieval with board-state matching to provide contextual scaffolding from both natural language queries and board topology.
We propose GraMRAG, a graph-memory guided multi-agent RAG framework that enables stable, multi-step multimodal reasoning by integrating a dynamic multimodal memory graph and a vision–text bridge.
We describe an semi-automated pipeline for automating attack graph construction for agentic pentesting, translating scanner outputs (Trivy, Semgrep, Nmap) into MulVAL predicates to enable neuro-symbolic vulnerability hunting with auditable reasoning.
We define the Interconnectedness Coefficient (IC), a semi-local graph measure that identifies connector vertices between cohesive network regions. IC favors weakly clustered focal nodes whose neighbors remain strongly clustered after removing the focal edge, capturing bridge-like roles.
We present a graph-based latent-retrieval approach for complete suffix prediction in recommendation. Prefixes and suffixes are encoded as directed attributed graphs using edge-conditioned graph neural networks, enabling joint modeling of event-level activities and durations, and latent retrieval of plausible suffixes.
COMPASS proposes steering distributed vector search with knowledge-graph-guided data placement and shard selection, preserving semantic locality beyond embedding-distance clustering. It detects communities in the knowledge graph, splits oversized communities, and uses factual relations to determine shard assignments and query routing, improving efficiency and relevance in scientific vector search.
GNN4PPM applies relational graph convolutional networks to predictive process monitoring, enabling multi-target predictions (next event, time to completion, and outcome) by leveraging rich relational context in event logs.
Cloud workflow scheduling uses a graph-attention driven hierarchical reinforcement learning framework to handle DAG-structured tasks. It assigns sub-deadlines to tasks to capture urgency and uses graph attention to guide decisions, balancing deadlines, resource usage, and energy.
HiGFRL proposes Hierarchical Graph Fusion-Driven Reinforcement Learning to schedule dependency-aware tasks in heterogeneous clouds. It captures high-order topological dependencies and tightly couples task and resource states to improve decision quality.
The paper explores physics-guided machine learning models to prescreen semiconductor point defects, predicting formation energies and zero-phonon lines to accelerate high-throughput defect screening prior to expensive DFT calculations.
We present IWC-Bench, an interactive benchmark for evaluating web application generation from a software testing perspective. It addresses limitations of static and purely interactive benchmarks by providing end-to-end evaluation across runtime behavior and testability.
We propose a spatially aware LLM-based multi-agent system for next-point of interest prediction, integrating spatial context such as geographic distance and neighborhood structure to overcome LLMs' spatial reasoning limitations and improve real-world predictive accuracy.
PCGNet proposes a unified model that jointly captures both general compatibility signals and individual user preferences for fashion matching, addressing the common decoupling assumption in prior work. By modeling how garment compatibility interacts with personal style and context, it aims to improve cross-selling recommendations in fashion.
This work analyzes the cross-stage decoupling between semantic signals used for item tokenization and collaborative signals used in the generation stage in generative recommender systems. It argues that current two-stage pipelines underutilize collaborative information during tokenization, causing misalignment with downstream generation, and proposes methods to better fuse signals across stages.
P3Rec introduces distillation of both prior and posterior preference reasoning from LLMs to improve recommendations, addressing the limitation of using a single-perspective distillation. Prior preferences capture users’ stable, long-term interests, while posterior preferences capture context-specific cues for the current decision, and the approach combines these to guide recommendations without heavy online LLM inference.
LazFormer scales industrial recommendation by scalable Transformer architectures with transferable generative pre-training, addressing the inefficiency of training a single ranking model to handle sparse and dense parameters. Pre-training initializes both sparse and dense components, enabling faster convergence and reduced resource consumption for downstream industrial recommendation tasks.
The paper proposes a multi-modal survival prediction framework that integrates clinical data, histology images, and genomics via a graph-guided mixture of experts. It leverages few-shot capabilities of large language models to handle cross-modal reasoning, enabling accurate tumor survival prognosis.
We introduce El Agente Potente, an agentic system that uses typed execution graphs plus a complementary coding mode to drive MLIP-powered atomistic simulations. This framework translates high-level scientific intent into adaptive, high-throughput campaign workflows while maintaining rigor and traceability.