←

Daily arXiv Papers

Graph Learning · LLM × Graph · Multi-Agent · Science

Showing 0 papers for 2026-09-19

🤗 Hugging Face daily top 5

most upvoted on 2026-09-18

DeepSeek-V4.1-Flash introduces a multimodal Mixture-of-Experts model with 552B parameters designed to push the limits of KV cache compression. It targets the compute, storage, and bandwidth bottlenecks caused by large KV caches in long-horizon agent workloads, aiming to reduce prefill costs and deployment expenses while supporting broader context.

DeepSeek· Hugging Face

MiniMax-H3 is an Omni-Modal Generative Model that unifies multimodal understanding with joint audio-visual generation in a shared latent space. The paper asks whether multimodal alignment can improve world reasoning and proposes a comprehensive evaluation framework organized around four complementary dimensions to assess the model's world understanding and cross-modal generation capabilities.

Length inflation in on-policy distillation arises from termination-token mismatch between base students and post-trained teachers. Across Qwen3, Llama, and Gemma, different EOS tokens attract stopping probability despite identical stopping sets, suppressing the student's chosen termination; the work shows that aligning termination behavior between student and teacher mitigates this issue.

Microsoft· Hugging Face ·GitHub ★6

SoL-Pi proposes recursively scaling auto-research loops at the harness layer to improve token efficiency for autonomous agents. It scales auto-research loops across many diverse environments for harness rollouts, producing reusable improvements that transfer beyond the development setting and move automated harness discovery toward production ready deployment.

NVIDIA· Hugging Face ·GitHub ★2264

An empirical study of harness design for coding agents investigates how individual harness components influence long horizon software engineering performance. Using a lightweight harness with a fixed execution loop, the study varies planning, action space and context management across four models and 176 settings on SWE-Bench Verified and Terminal-Bench 2.1, revealing component-level effects on performance.

Zoom Communications· Hugging Face