←

Daily arXiv Papers

Graph Learning · LLM × Graph · Multi-Agent · Science

Showing 0 papers for 2026-09-26

🤗 Hugging Face daily top 5

most upvoted on 2026-09-25

We introduce WROP, a data infrastructure of 150 hand-designed cognitive-science-inspired tasks across six categories, designed to train world models to acquire object permanence. The goal is to endow video-based world models with human-like physical priors and robust reasoning about occlusion and solidity.

Carnegie Mellon University· Hugging Face ·GitHub ★12

This work provides evidence that Transformers exhibit a Superposition Linearity: when inputs from distinct text streams are linearly combined, the model's next-token distribution is roughly a superposition of the originals. The linearity appears to be an intrinsic property of the Transformer architecture and tends to diminish with ongoing pretraining, and the authors also show that linearity can be induced or controlled.

Hugging Face

WanPE introduces a 397B parameter prompt-enhancement model trained on 1.05 million real-world videos to master director-level cinematic planning for text-to-video generation. It formulates shot-level plans via video-grounded reverse construction to guide actions, camera trajectories, lighting, and sequencing.

Hugging Face

OmniEchoBench provides a unified benchmark for spatial audio-visual perception and audio-vision-language navigation in embodied agents. It includes six tasks over 197 real-world spatial audio-visual scenes, 2,972 question-answer pairs, and 900 navigation samples using first-order ambisonics (FOA).

PKU-VaLuE-Lab· Hugging Face ·GitHub ★11

Agent-Editing World Model (AEWM) proposes rethinking world-model-based agents by enabling edits to the agent's internal world model to remove execution-dependent, outdated plans, reducing task-state contamination and improving long-horizon decision making.

Renmin University of China· Hugging Face ·GitHub ★4