Graph Learning · LLM × Graph · Multi-Agent · Science
Showing 2 papers for 2026-09-04
Nothing cleared the bar today. Papers read in full today: 0.
Autonomous agents currently lack operational knowledge that makes methods work in the real world, a layer encapsulated in GitHub repos and papers but too large to load during tasks. Repo-To-Skill proposes distilling this know-how into compact, verified AI4AI skills that can be loaded and reused by agents during execution. The paper outlines a framework to extract, verify, and compress repository-level know-how into usable skills.
HarnessDev shifts evaluation from task outputs to the runnable agent harness itself. It presents a benchmark to study whether LLMs can design and evolve their own execution infrastructure, not just perform tasks, with two stages capturing harness creation and subsequent evolution under varying conditions.
Aspire investigates self-evolution driven by vague goals like 'become a better physicist.' Unlike prior work that starts from explicit tasks, Aspire offers a vague-goal-driven self-evolution benchmark where models must interpret goals, identify capability gaps, plan learning, and assess progress using natural language prompts.
SolarWM introduces an open, end-to-end foundation for long-horizon video world models, from data preparation to inference. It tackles data heterogeneity (temporal scales, camera geometry, quality, captions) and diverse video backbones by proposing a reconfigurable, multi-source training approach that aligns supervision across sources and models to enable scalable, reproducible results.
EarlyEval proposes out-of-task efficiency by predicting a task’s final outcome early, allowing cheaper agent evaluation. By forecasting results during task execution, it enables early termination or reduced computation, complementing distillation-based benchmarks to lower per-task costs in iterative agent development.
GRASP introduces a Graph-Retrieval Automated Scoring Pipeline for label-free, multi-topic short-answer science exams. It handles responses where answers to several distinct topics are merged into a single paragraph without markup, by using graph-based retrieval to identify topic-related segments and assign scores across topics. The approach enables label-free, multi-topic scoring and is validated on relevant datasets.
R2Adapter proposes a routing and rewriting adapter to make Hybrid Retrieval-Augmented Generation (RAG) more efficient. It dynamically routes and rewrites user queries to combine text-based and graph-based retrieval paths, addressing simple and complex (relational or multi-hop) reasoning with lower latency than fixed strategies. The method reduces inference cost while preserving or improving reasoning capabilities in hybrid RAG systems.