Graph Learning · LLM × Graph · Multi-Agent · Science
Showing 31 papers for 2026-10-01
Nothing cleared the bar today. Papers read in full today: 14.
Raven proposes a framework called Raven, the Harness of Harnesses, to autonomously construct, improve, and orchestrate specialized ‘harnesses’ for agentic intelligence across domains. It addresses the scalability and generality problems that arise when harness complexity grows and domain-tied solutions proliferate, by enabling composition and autonomous harness evolution. The goal is long-horizon, cross-domain AI workflows that can re-use and adapt harnesses without manual re-engineering.
MaLiang-Harness defines a unified framework that treats image/video generation as a running visual program driven by MLLMs. It aims to close the Program-to-Visual (P2V) gap by organizing generation into construction, inspection, and revision, while keeping the evolving program, its history, and its verification as first-class artifacts. This enables explicit control, error checking, and iterative refinement to align output with requested composition, appearance, and motion.
This survey reviews in-context learning (ICL) for robots, focusing on how demonstrations and interaction steer fixed neural policies toward new tasks. It organizes methods around four interface families that map contextual evidence to execution: context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- or agent-based execution. The review compares approaches, clarifies trade-offs, and highlights open challenges and applications in robotics.
Omni-IO Skills introduces a plug-and-play Agent Harness that makes existing agents omni-native by organizing capabilities into hierarchical Skills and a standardized multimodal interface. It addresses fragmentation across modalities (text, vision, audio, video, documents, 3D assets, code) and the orchestration of procedures, dependencies, intermediate assets, and cross-turn revisions. The result is a unified, extensible framework for building general-purpose agents that can operate across domains without bespoke re-engineering.
This work studies how on-policy distillation (OPD) scales across different teacher-student configurations in large language model (LLM)–reinforcement learning setups. By examining weak-to-strong, same-base, and strong-to-weak pairs, it shows that early OPD dynamics enter a regular, useful-transfer phase, during which the gold score improves roughly with the distance between the student and reference policies (measured by KL divergence). The results clarify how capability transfers grow with model scale and teacher strength, offering guidelines for scaling OPD programs.
CARAT asks whether materials LLMs truly reason about crystal structures or merely recite training data. It introduces eight matched views, GraphSpace, and techniques like evidence injection and paired inference to separate reasoning from memorization.
Stable Transformers for Graph Generation analyzes why deeper Graph Transformers can underperform in generation tasks due to contraction of representations under repeated self-attention. It studies the denoiser's spectral dynamics and proposes stability-enhancing strategies.
The paper proposes a DAG-based planning framework for deep research agents, organizing sub-tasks as a dependency graph to enable parallel execution and isolation. It argues that the common plan-then-patch approach—planning upfront and only repairing after failures—is brittle for deep research tasks. To address this, it introduces evaluate-then-grow: the agent continually evaluates evidence and grows or repairs the DAG during execution, improving robustness and adaptability.
Residual Trajectory Distillation for Generative Retrieval shows that in generative retrieval with residual-quantized codes, standard supervision only trains the selected codes and discards residual trajectories that carry information about subsequent quantization decisions. The authors propose distilling these residual trajectories to preserve richer decoding information, improving retrieval accuracy and robustness.
Warm-starting PDE solvers with any-dimensional machine learning derives mathematical conditions under which a solver trained in small dimensions can be applied in zero-shot to higher dimensions, leveraging symmetries in the PDE and initial data.
TAGGRAPH introduces a controlled evaluation framework for agent persistent memories using tag-augmented graphs, comparing localized configurations, chronological edges, and diffusion-based retrieval against BM25 and OpenClaw.
EB-GAD proposes training-free node-level graph anomaly scoring by casting anomaly scoring as a finite-horizon control problem and calibrating scores with empirical Bayes. This approach reduces the entanglement between graph trust, spectral weighting, and score choice, yielding more stable and interpretable anomaly scores across graph regimes.
GraphMAS offers a systematic benchmark of multi-agent coordination for graph learning with LLM-based agents. It evaluates how coordination, evidence aggregation, and adaptive control affect graph reasoning across diverse tasks, comparing single-agent and multi-agent setups.
GraphCert bootstraps agentic graph reasoning with certified evidence rubrics, reducing reliance on large QA data and external LLM services. Rubrics guide multi-step reasoning and anchor claims with verifiable evidence.
Scalable Approximate Algorithm for Dynamic Densest Subhypergraphs with Solution-Guided Maintenance presents CAP, a dynamic algorithm that maintains a (1+epsilon)-approximate densest subhypergraph under hyperedge insertions and deletions. The approach keeps a solution candidate plus endpoint allocations to bound the optimum density, enabling scalable updates with theoretical guarantees.
When the Label Ignores the Request: Auditing Policy-Selected Targets in Synthetic Conversational Music Recommendation investigates how synthetic dialogue labels align with user requests. The study audits exactly-named-song scenarios in the RecSys Challenge 2026 TalkPlay benchmark to reveal gaps between policy-selected labels and user intent and discusses implications for evaluation.
Life-Bench is a fully synthetic, human-verified multimodal benchmark with over 11,800 QA pairs across 10 tasks to evaluate multimodal personalization beyond simple concept recognition.
When LLM-Inferred User Context Adds Value in Production Streaming Recommendation evaluates when semantic profiles derived from large language models improve production streaming recommender performance over traditional aggregate embeddings. The study analyzes conditions across domains where unstructured interaction histories benefit from LLM-generated context, and discusses the tradeoffs in computational cost and latency.
NodeGround provides a node classification benchmark in the era of graph foundation models by evaluating GFMs on 51 datasets against dataset-specific supervised learners under two label-availability regimes. It standardizes partitions and model selection to compare accuracy and computational cost.
CollabFlow presents recursive self-improvement for collaborating agents, enabling self and cross-task refinement to overcome fixed collaboration topologies and error propagation.
World-as-Graph (WAG) presents a graph-based, object-centric world model that encodes relational inductive biases to capture object–object interactions and temporal structure. The model uses latent-space graphs to represent objects and their relations and to predict future dynamics.
About the Influence of Workflow Topology on Task Intensity Prediction through Graph Learning investigates how DAG topology affects task-level resource intensity predictions. It shows that topology-aware features improve accuracy and aid resource provisioning.
EvoSteer proposes online self-evolving graph orchestration to continually adapt agent teams, addressing post-hoc evolution, credit diffusion, and skill admission with a reference-anchored credit assignment.
ChronoGraph defines functional 4D scene graphs that connect actions to semantic and geometric state changes over time, enabling interaction understanding and grounded planning.
AdaM-Rec introduces Adaptive Modality Routing for multimodal recommendation, dynamically adjusting the emphasis on visual and textual signals by query context. Static modality fusion assumes fixed importance, which can misfire when some requests require fine-grained visuals while others rely on semantic text. The method routes requests to modality-specific components or adjusts fusion weights to improve accuracy and efficiency.
Routing Between Generative and Collaborative User Profiles: A Serving-Time Gate for Controllable Novelty proposes a serving-time routing gate that assigns each user to either a collaborative sequential recommender or a recommendation model driven by an LLM-generated user profile. This enables explicit control over novelty versus stability while balancing serving cost. Experiments on a real streaming dataset show improved tradeoffs between recommendation quality and latency.
Decision-Oriented Recommendation Reranking: An Empirical Study of Jev conducts an empirical comparison between decision-oriented LLM-based reranking (Jev) and traditional reranking approaches. It analyzes the tradeoffs between recommendation quality and serving efficiency and identifies conditions under which Jev provides advantages for structured decision tasks.
RetroGEF is a flow-based model for single-step retrosynthesis that starts from the target molecule and performs dynamic graph edits to construct reactant graphs. It models connectivities and graph sizes changes, including the introduction of new reactant components, in a single forward pass.
Autoregressive Frontier Expansion introduces an autoregressive framework to grow tree-like structures with graph machine learning. It expands the generation frontier to produce realistic branching morphologies for 3D data when real-world samples are scarce.
Riemannian Flow Models with Reinforcement Learning for Molecular Crystal Structure Prediction (CG-OMatG) presents an equivariant, flow-based generative model for molecular crystals. It leverages symmetry-aware flows and RL to improve structure predictions amid polymorphism and complex packing.
MANET-GNN develops a learned decentralized optimization framework for power allocation in dynamic, multi-channel MANETs. It formulates end-to-end throughput maximization under network constraints and enables decentralized control for various traffic modes.
Exploring Forum Post Retrieval with Generative Modeling investigates applying generative recommendation to Facebook Forum posts. Because Forum data are sparse, the work uses transfer learning along two axes: training on a broader corpus of Facebook Groups engagements and reusing hierarchical prefix-based modeling to adapt to Forum. The study demonstrates gains in retrieval quality on Forum data.
Generative End-to-end Ad Retrieval at Douyin identifies two bottlenecks in scalable generative retrieval: representation collapse under distribution shifts and item collisions in a large candidate pool. The paper analyzes these bottlenecks and proposes methods to stabilize end-to-end training and improve token-based retrieval quality at industrial scale.
MASCRDM introduces a multi-agent system to detect and mitigate compliance risks during LLM training, addressing the limitations of static detection and filtering through real-time coordination and remediation.
RankEvolve presents a reliable auto-research harness for evolving ranking models, featuring an Executable Operating Protocol that ensures reproducibility, prevents data leakage, and tracks changes across iterations.
KUAISHOU Explorer LLM-Rec Challenge 2026: Reasoning Generative Recommendation reports on the Challenge focusing on reasoning in generative recommendation. Building on OneRec/OneRec-V2 semantic ID models, the team demonstrates deployment potential and scalability of autoregressive next-item generation in industrial recommender systems.