Graph Learning · LLM × Graph · Multi-Agent · Science
Showing 0 papers for 2026-10-03
OneStreamer unifies perception, memory, and proactive response for streaming video LLMs. It learns query-independent evidence recording and task responses through a shared proactive generation process, and introduces Proactive Hierarchical Caption Memory (PHCM) that produces time-grounded local-detail captions and summaries of completed events. This memory enables reusable factual evidence while preserving real-time perception during streaming.
This paper systematically studies distillation dynamics by focusing on rollout policy in a controlled setup. They vary rollout policy, token-level KL direction, and learning rate across Llama3 and Qwen2.5, covering scientific, medical, and arithmetic reasoning tasks to isolate the effects of on-policy vs off-policy distillation on forgetting, update sparsity, and generalisation.
They propose Adaptive Reward Routing (ARR) to dynamically optimize multiple rewards for joint audio-video diffusion using forward-process RL. The method addresses when and how to apply reward-driven updates and how to combine competing rewards, which prior work fixes. ARR continuously adapts routing locations and reward weights during training to align with evolving model functions, improving modality quality, semantic alignment, and temporal synchronization.
They introduce PoS, an inference-time framework that constructs and maintains explicit belief states as the agent's decision context for long-horizon tasks. Each belief encodes an estimate of the current world state together with unresolved task requirements, making explicit what the agent still needs to learn and accomplish. PoS also validates consistency and monitors task progress to keep beliefs reliable and actionable.
They propose Hierarchical Continuous Diffusion Language Models to overcome limitations of parallel decoding in discrete diffusion LMs and the mismatch in continuous diffusion where the denoiser only sees a shared state. The hierarchical approach ties denoising dynamics to valid token configurations by introducing hierarchical structure that preserves token dependencies during diffusion, enabling more coherent bidirectional reasoning and global constraint satisfaction.