←

Daily arXiv Papers

Graph Learning · LLM × Graph · Multi-Agent · Science

Showing 20 papers for 2026-09-03

★ Must read

full paper read, not just the abstract
★ MUST READ · high confidence
A Power Law in Logarithm's Clothing: On the Scalability of Graph-Based Vector Search
Graph Theory
University of Toronto

The load-bearing contribution is a falsifiable reframing of ANN scalability: holding recall fixed, the expensive part is not reaching the query vicinity but exploring a locally crowded neighborhood, whose size scales as k(1+ε)^d_int. The paper does more than report a benchmark win: across eight datasets, 90–99% recall, query-hardness strata, and HNSW/Vamana build settings, it finds near-linear log–log curves, while SpaceV shows the predicted flattening at larger scale; the theory usefully links those observations to LID growth and gives configuration-sensitive cost models. This is a strong diagnostic result that should alter how vector-search users forecast capacity and interpret advertised logarithmic complexity.

Against prior work

This must be measured against HNSW (Malkov and Yashunin, 2018), Vamana/DiskANN (Subramanya et al., 2019), and the exact proximity-graph theory behind Randomized Neighborhood Graph/SNG/MRNG and Indyk & Xu (2023). Before this paper, practitioners had HNSW/Vamana benchmarks at isolated corpus sizes and a widely repeated, weakly supported logarithmic-scaling folklore; this paper supplies fixed-recall scaling measurements across sizes and argues that the practically relevant finite-scale regime is instead dc ∝ N^c, driven by growing local intrinsic dimensionality, before a later slowdown. The ingredients—LID-based hardness, local volume growth, and curse-of-dimensionality arguments—are established, but the specific fixed-recall scaling diagnosis and its connection to graph beam-search cost do not appear to be a mere renaming of a known result.

The strongest asymptotic claims are conditional—local uniformity, in-distribution queries, Euclidean geometry, and the assumption that fixed recall corresponds to exploring a multiplicative-radius neighborhood—and the practical evidence is limited to HNSW and Vamana, with billion-scale confirmation for one Vamana configuration rather than production variants, updates, filtering, or adversarial workloads.

◇ Potentially interesting

read in full, rated just below must-read — your call
University of Maryland, Baltimore County

The load-bearing contribution is the empirical reframing of triple F1 as a function of a matching predicate and document-scoping rule, backed by rescoring ten cached system outputs under eight protocols and by a mechanical-matcher calibration test. The full text does deliver the central diagnostic result: rankings are matcher-sensitive, pooled matching inflates true positives by 6.4%, and the validation ablation usefully finds that hand-written rules improve precision on all tested hosted configurations but harm it on all tested local ones. This is a strong, unusually candid audit, but it is a well-executed domain-specific evaluation study rather than a broadly new graph-learning mechanism or theory.

The strongest calibration evidence is out-of-domain and reuses GRID's reported LLM-judge agreement rather than independently adjudicating CTI triples; moreover, the hosted/local validation split confounds backbone, prompting, decoding, and serving stack, so it cannot establish the proposed schema-mismatch mechanism causally.

🤗 Hugging Face daily top 5

most upvoted on 2026-09-02

StudentSim is a training framework for building LLM-based student simulators from sparse per-student data. It provides a scalable surrogate of student behavior that can guide tutoring policy and guidance selection. By integrating strengths of state-tracking models and guidance-aware LLMs, it aims to enable more effective, individualized tutoring without costly real learners.

Microsoft Research· Hugging Face ·GitHub ★10

Qwen-Drive-1.0 is an initial step toward a vision-language foundation model for autonomous driving. It preserves the pretrained VLM architecture and adds 3D perception, VQA, and motion planning in a unified framework. A BEV perception head jointly performs 3D object detection, semantic occupancy prediction, and BEV map segmentation, providing an explicit interface to 3D scene structure that can be probed from shared representations.

Qwen· Hugging Face

SMELT studies compute-matched looping in Mixture-of-Experts transformers. It keeps per-token FLOPs, total non-embedding parameters, and KV cache fixed while looping a shared block, then experiments to derive a recipe. The result, SMELT (Sparse MoE Transformer, middle layers Loop Twice), loops the middle half of the layers twice and matches the Baseline budgets, scaling up to 54B non-embedding parameters across four sizes.

ByteDance Seed· Hugging Face

UI-Venus-2 presents a general-purpose foundation GUI agent for mobile, web, and desktop, operating under a unified closed-loop reasoning-action framework. The work targets real-world deployment by scaling three dimensions: environment coverage, data quality/scale, and task reliability. This aims to bridge the gap between benchmark progress and dependable, real-world digital task automation.

Ant Group· Hugging Face

H3-World turns language understanding into world control. It shows that as large video generators become capable, language is a natural control interface, with MiniMax-H3 enabling zero-shot control of character behavior and camera motion. H3-World then turns this coarse language interface into precise, temporally grounded world control without dedicated action modules, representing each action as a struct.

All papers

18
Oracle, will I ever learn? A study of prediction convergence and complementarity across link prediction models
Knowledge Graph Graph Learning

Investigates prediction convergence and complementarity across link prediction models. The study shows that different models and even different training runs yield substantially different predictions for the same query, suggesting that models capture complementary signals. The paper discusses implications for ensemble methods and reliability of link-prediction benchmarks.

Seed-Anchored Budget-Bounded Graph Rendering for Question Answering on Industry-Standard Power-Grid Information and Exchange Models
Graph Learning Graph Theory

Introduces seed-anchored graph rendering to question answering over industry-standard power-grid CIM models, enforcing a fixed context budget. It greedily renders seed-local evidence within a hop bound to guarantee preservation of answer-bearing units, offering a checkable condition for correctness.

Omega-N: Interpretable Structural Node Descriptors and Their Applicability Domain
Graph Theory

Introduces Omega-N, a family of interpretable structural node descriptors derived from A^3-based measures. By using diag(A^3), the approach captures non-redundant structural information and enables node-wise attributions, with analysis of the applicability domain relative to spectral baselines.

Higher-order rich clubs and configuration models on general directed hypergraphs
Graph Theory

Proposes higher-order rich clubs and configuration models for general directed hypergraphs, enabling analysis of higher-order interconnections beyond pairwise edges. The hyper-rich club pipeline identifies whether high-centrality nodes are more densely interconnected than expected by chance, using configuration models for directed hypergraphs.

Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems
Multi-Agent Graph Learning

Introduces Codebook Agent for amortized topology design in LLM-based multi-agent systems. Rather than sampling from the full adjacency space, it uses a fixed codebook of candidate topologies that survive reward filtering, yielding a small set of effective topologies and more stable performance.

Spectral Initialization and Scheduled Graph Smoothness for Uncertain Knowledge Graph Completion
Knowledge Graph Graph Learning

Introduces QUEST, a UKG completion method that adds no trainable parameters to the standard confidence-distribution learning pipeline. It initializes entity embeddings from the confidence distribution using a spectral-like initialization and applies a scheduled graph-smoothness process to propagate confidence across the graph. This parameter-free approach preserves global community structure and improves pseudo-label quality for uncertain triples.

A Comparative Study of Graph Representations for GNN-Based Power Grid Control in L2RPN
GNN Graph Learning

Conducts a controlled comparative study of several graph representations for GNN-based power-grid control in the L2RPN benchmark. The representations include physical topology, electrical-sensitivity, and hybrid variants. The results show that matching graph complexity to task granularity matters more than maximizing representational richness, highlighting the value of controlled representation choice.

KGVoyager: Knowledge Graph Agnostic Question Answering via Agentic Navigation
Knowledge Graph Multi-Agent

Introduces KGVoyager, a KG-agnostic, agentic QA system that generates SPARQL from natural language by dynamically discovering graph structure and semantics. Using a think-act-observe loop with search, exploration, and execution tools, it maps terms to graph IRIs and identifies structure, enabling effective KGQA without relying on ontologies.

Import What You Need: Learning When and How to Augment EHR Graphs with External Knowledge
Graph × Science Knowledge Graph

Proposes ReTA, a reinforcement-learning-based dynamic topology augmentation framework for EHR graphs. It treats knowledge graph import as a per-visit, budget-aware decision problem, selecting external KG nodes and edges according to the patient's evolving state. The goal is to alleviate sparsity and irregularity in longitudinal EHR predictions by making topology augmentation conditional on patient state.

C$^{3}$T: Counterfactual Causal Reasoning for Sentiment Shifts in Social-Media Conversation Trees
Graph Theory

Defines C3T, a counterfactual causal framework to analyze sentiment shifts in rumor-centered social media conversation trees. It treats discourse moves as interventions and asks (i) what sentiment a reply expresses, (ii) whether it shifts from its parent, and (iii) which prior message most plausibly drove the sentiment; uses counterfactual reasoning to explain dynamics.

Refining Heuristic-Based Bitcoin Address Clustering with Graph Neural Networks
GNN Graph Learning

Refines heuristic-based Bitcoin address clustering by grounding the clusters with contrastive embeddings produced by graph neural networks. The approach learns GNN-based representations to refine heuristic clusters and uses contrastive objectives to improve modularity while reducing merge errors across distinct users. Experiments demonstrate improved clustering quality and robustness.

PEARL: Path-Entity Aligned Relational Learning with Contextual Subgraphs for Inductive Knowledge Graph Completion
Knowledge Graph GNN Graph Learning

Proposes PEARL, Path-Entity Aligned Relational Learning for inductive knowledge graph completion. It models paths as context-conditioned reasoning signals by constructing query-contextual subgraphs that align paths with the surrounding context, enabling transfer to unseen entities.

GRAND-HC: Graph-Refined Author Name Disambiguation
Graph Learning

Presents GRAND-HC, an end-to-end graph-refined author name disambiguation framework that mitigates long-tailed author biases and unreliable cluster counts. It builds a heterogeneous paper graph incorporating co-authorship, affiliation, and venues to jointly disambiguate authors in a complete pipeline.

Dual-Metric Partitioning with Adaptive Kernel Execution for Efficient GCN Acceleration
GNN

Proposes DualGCN, a GPU-accelerated GCN framework that addresses width and depth workload imbalances with dual-metric graph partitioning and adaptive kernel execution. The approach reduces irregular memory access and improves throughput for large graphs.

Hybrid Retrieval-Augmented Generation with Knowledge Graph Expansion, RRF Fusion, and Per-Chunk Grounded Evaluation for Enterprise Document Search
GraphRAG Knowledge Graph

Proposes DocuSearch, a hybrid retrieval-augmented generation system for enterprise document search. It combines knowledge graph expansion, reciprocal rank fusion, and per-chunk grounded evaluation to provide accurate and verifiable answers over large corporate repositories, demonstrated in a telecom network-operations setting.

SSAKG 2.0: An Open-Source Package for Structural Associative Sequence Memory and Context-Based Retrieval
Knowledge Graph Graph Learning

Presents SSAKG 2.0, an open-source package for Structural Sequential Associative Knowledge Graphs used as sparse memory to reconstruct sequences from partial contexts. It introduces memory-bit-level search algorithms to efficiently traverse graph connections.

Semantics-Guided Automatic Tensorization for Multiobjective Evolutionary Algorithms: A Multi-Agent Framework
Multi-Agent

Proposes Semantics-Guided Automatic Tensorization for MOEAs, a framework to restructure multiobjective evolutionary algorithms for tensor-enabled hardware without changing the underlying optimization mechanism. It treats tensorization as semantics-guided computation rearrangement to exploit modern accelerators.

MGDiff: Multi-Interest Sequence Recommendation with Masking GNN-Guided Diffusion
GNN Generative Rec

Introduces MGDiff, a multi-interest sequence recommender using Masking GNN-Guided Diffusion. It uses a semantic guidance framework to extract item semantics and separate user intents, and a weight-adaptive masking GNN to reconstruct missing user–item links during diffusion, improving accuracy and reducing bias.