Graph Learning · LLM × Graph · Multi-Agent · Science
Showing 0 papers for 2026-02-14
The paper argues that a society of self-evolving LLM agents cannot simultaneously achieve continuous self-improvement, complete isolation, and safety invariance. Using an information-theoretic framework that formalizes safety as a divergence measure, they prove that such a closed-loop system is impossible, with empirical evidence supporting the claim. The result highlights an unavoidable safety-trade-off in self-evolving AI societies.
The paper introduces Composition-RL to better utilize limited verifiable prompts for RL with verifiable rewards (RLVR). It addresses the problem that focusing only on hard prompts (pass rate 0) wastes data as training progresses, since easy prompts (pass rate 1) become prevalent. Composition-RL proposes composing and prioritizing prompts to maintain informative coverage and improve data efficiency in verifiable prompt learning.
DeepGen 1.0 is a lightweight 5B unified multimodal model for image generation and editing that competes with larger models in capability while reducing training and deployment costs. To overcome compact-model limitations in semantic understanding and fine-grained control, it introduces Stacked Channel Bridging (SCB), a deep alignment framework that fuses hierarchical features from multiple vision-language model layers to improve performance.
This work analyzes On-Policy Distillation (OPD) and shows it is a special case of dense KL-constrained reinforcement learning where reward and KL weights are equal. It then introduces Generalized On-Policy Distillation (G-OPD), extending OPD with reward extrapolation and broader settings to enhance student learning under varied reward structures.
GigaBrain-0.5M* introduces a Vision-Language-Action model learned from world-model-based reinforcement learning. Unlike direct multi-step action prediction, it leverages video world models pre-trained on large-scale videos to improve spatiotemporal reasoning and future prediction. Built on GigaBrain-0.5 with 10k+ hours of robotic video data, it demonstrates improved VLA learning.