News Β· Leaderboard

Latest news, beside the leaderboard

Read what just changed, and see who is ahead. Those are the two doors on this page.

Plans
X
XunzhuoXunzhuoLiuΒ·1h ago
πŸš€ Release

vLLM Semantic Router open-sources Decision 3.0: six multimodal decision models from 0.6B to 27B, #1 on Jev Decision Index in its own evaluation

On Oct 11 the vLLM Semantic Router (vllm-sr) team released the Decision 3.0 family (d3, d3-flash, d3-mini, d3-nano, d3-lite, d3-edge) on Hugging Face under Apache-2.0. Built on Qwen3.5/3.8, the models take text, images and video and return probabilities for choice, yes/no and score questions in one call without generating text. Vendor-reported: the 27B d3 scores 64.7 on Jev Decision Index 0.3.1, ahead of Perplexity Decider v1.1 (62.8) and Decision 2.0 (55.9), with a 55 ms median text latency on one AMD MI325X.

vLLM Semantic Router open-sources Decision 3.0: six multimodal decision models from 0.6B to 27B, #1 on Jev Decision Index in its own evaluation
⚑ Key Takeaways
  • β€’Six sizes: d3-edge 0.59B, d3-lite 0.85B, d3-nano 2.2B, d3-mini 4.5B, d3-flash 8.4B, d3 26.1B, all Apache-2.0 (HF collection)
  • β€’Vendor-reported Jev Decision Index 0.3.1: d3 64.7 vs Perplexity Decider v1.1 62.8, Jev 60.1, Decision 2.0 55.9
  • β€’Vision Index 0.3.1: d3 71.6 vs Perplexity Decider v1.1 70.6 and JEV-27B-VL 69.6
  • β€’Median latency on one AMD MI325X, batch 1: d3 55 ms text / 273 ms image / 553 ms 10-second video; d3-edge 6.7 ms text
  • β€’No text generation: every option gets a probability, and several Choice / Yes-No / Score questions share one call
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·1h ago
πŸ”₯ Trending

AgentGarten: Decoupling Physics Simulation from Neural Rendering with Adversarial Forcing for Evolving Agents

Interactive virtual environments allow embodied AI agents to learn through exploration and practice, but creating worlds that are simultaneously faithful (rigid state dynamics and programmatic rules) and realistic (photorealistic visual distributions) has remained a critical bottleneck: physics engines lack visual fidelity, whereas generative video models suffer from compounding visual hallucinations during extended agent interactions. MirroS Lab and researchers introduce AgentGarten (arXiv:2610.12374, repo: github.com/MirroS-Lab/AgentGarten, site: mirros-lab.github.io/agent-garten), an open-source framework that decouples physics simulation backends from a shared neural renderer. Simulators maintain ground-truth state and execute code-defined rules without ingesting rendered frames, strictly preventing error accumulation. Visual observations are synthesized in real-time (>30 FPS) from exported geometric conditions via a neural renderer trained with Adversarial Forcing, making history prefilling differentiable through exact replay. Agents learn through visual perception, distilling execution trajectories into playbooks that succeeding generations inherit. In benchmarks such as hide-and-seek, agents autonomously discover obstacle-moving and ramp-climbing tactics in just 4 to 10 rounds (compared to tens of millions of rounds in classic RL), with code and environments fully open-sourced.

AgentGarten: Decoupling Physics Simulation from Neural Rendering with Adversarial Forcing for Evolving Agents
⚑ Key Takeaways
  • β€’MirroS Lab releases AgentGarten, decoupling physics simulation backends from neural rendering to eliminate compounding hallucinations
  • β€’Introduces Adversarial Forcing differentiable replay, clocking 37.4 FPS real-time neural rendering on a single NVIDIA H100 GPU
  • β€’Emerges complex multi-agent collaborative strategies in just 4 rounds, outperforming classic RL sample efficiency by orders of magnitude
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·1h ago
πŸ”₯ Trending

TokenRouter: Tsinghua Unveils NeurIPS 2026 Serving System for Token-Level LLM Routing, Boosting Throughput by up to 64x

Token-level LLM routingβ€”routing easy tokens to small models and switching to frontier models for critical reasoning tokensβ€”presents an optimal trade-off between output quality and inference cost. However, current LLM serving engines (such as vLLM and TGI) are architected for single-model execution, suffering severe step desynchronization, batch admission latency, and extreme orchestration complexity when handling high-frequency token transfers across heterogeneous models. Researchers from Tsinghua University's NICS Lab present TokenRouter (arXiv:2610.12242, accepted at NeurIPS 2026, code: github.com/thu-nics/TokenRouter), an efficient and developer-friendly serving system engineered for token-level routed LLM inference. TokenRouter introduces a 'request-centric programming and model-centric execution' paradigm, enabling developers to write routing logic intuitively per request while the runtime handles asynchronous execution across subservers. Incorporating a Delayed-Batching Scheduler parameterized via discrete-time Markov chain modeling, a decoupled tri-loop architecture, and an instant handoff-resume mechanism, TokenRouter achieves 2.01x to 64.15x higher decoding throughput compared to existing serving platforms across diverse routing workloads, with complete code open-sourced.

TokenRouter: Tsinghua Unveils NeurIPS 2026 Serving System for Token-Level LLM Routing, Boosting Throughput by up to 64x
⚑ Key Takeaways
  • β€’Tsinghua University open-sources TokenRouter (NeurIPS 2026), an optimized serving architecture for token-level LLM routing
  • β€’Introduces request-centric programming and a delayed-batching scheduler, delivering 2.01x to 64.15x higher decoding throughput
  • β€’Decoupled tri-loop execution cuts P99 inter-token latency by 43.6%, reducing enterprise cloud GPU costs by 50% to 70%
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·1h ago
πŸ”₯ Trending

LoGRA: Scaling LLM Reinforcement Learning via Low-Rank Gradient Sketches, Slashing Memory by 45.7% for 27B Models

Reinforcement learning (RL) post-training is foundational to scaling reasoning capabilities in large language models. However, its memory footprint remains a prohibitive bottleneck: storing full-sized gradient buffers alongside multi-state Adam optimizer moments frequently triggers out-of-memory (OOM) failures when attempting to train large models (such as 27B parameters) on standard compute hardware. In newly released research (arXiv:2610.06647), researchers introduce LoGRA, an efficient RL post-training framework available within the open-source Molt library (github.com/skzhang1/labs-molt). LoGRA compresses training representations by retaining useful learning signals inside compact low-rank gradient sketches accumulated directly during backpropagation, facilitating both memory-efficient parameter updates and streamlined policy synchronization. To counteract instability from gradient compression, LoGRA introduces a Predicted-KL step control mechanism that forecasts policy change magnitudes prior to parameter updates, dynamically attenuating disruptive updates. Evaluated across reasoning benchmarks, LoGRA reduces average training memory by up to 45.7% with zero degradation in task performance, successfully sustaining stable RL training of a 27B model for over 1,100 steps on a single eight-GPU node where standard dense Adam crashes with OOM.

LoGRA: Scaling LLM Reinforcement Learning via Low-Rank Gradient Sketches, Slashing Memory by 45.7% for 27B Models
⚑ Key Takeaways
  • β€’LoGRA framework released, introducing low-rank gradient sketches to slash LLM RL training memory by up to 45.7%
  • β€’Breaks single-node 8-GPU memory barriers, sustaining stable RL training of a 27B model for over 1,100 steps
  • β€’Predicted-KL step control prevents policy drift, achieving full parity with standard AdamW on reasoning benchmarks
Read details→
ADSponsored
v
vLLMvllm_projectΒ·3h ago
πŸš€ Release

vLLM adds NVIDIA Vera Rubin NVL72 support: up to 7.84x per-GPU throughput over GB200 on AgentX with MiniMax M3

On Oct 9 the vLLM team, with Inferact, Red Hat and NVIDIA, shipped early Vera Rubin NVL72 support: CUDA 13.4 nightly images already serve DeepSeek, Kimi, GLM and MiniMax. Vendor-reported results: up to 7.84x per-GPU throughput vs GB200 on SemiAnalysis AgentX (MiniMax M3, matched interactivity) and up to 3.7x vs GB300 NVL72 on the MLPerf Inference v6.1 VLM benchmark (Qwen3-VL-235B-A22B).

vLLM adds NVIDIA Vera Rubin NVL72 support: up to 7.84x per-GPU throughput over GB200 on AgentX with MiniMax M3
⚑ Key Takeaways
  • β€’Versus GB200 NVL72: 5x NVFP4 FLOPS, ~2.4x HBM4 bandwidth, 1.7x bidirectional NVLink, 2–4x faster softmax exponentials (vLLM blog)
  • β€’AgentX with MiniMax M3: up to 7.84x per-GPU throughput at matched interactivity, 5.18x under a 150 TPS cap (vendor-reported, early)
  • β€’MLPerf Inference v6.1: Qwen3-VL-235B-A22B on vLLM + Dynamo, up to 3.7x vs GB300 NVL72
  • β€’CUDA 13.4 locality domains split MoE weights for ~1.2x average MoE-layer speedup in small-token decode
  • β€’Image: vllm/vllm-openai:cu134-nightly (CUDA 13.4 + PyTorch 2.15)
Read details→
w
wheresryan22wheresryan22Β·5h ago
πŸ”₯ Trending

Anatomy: an open-source Claude Code skill that turns technical ideas into interactive isometric machine figures

Anatomy (MIT) is an Agent Skill that has Claude Code invent a physical machine whose mechanism embodies a concept, then build it as an interactive isometric SVG/WebGL/3D (beta) figure with its own geometry audit and headless-Chrome self-checks. Created Oct 8, v1.0.0 shipped Oct 10, ~1,970 stars and 150 forks so far.

Anatomy: an open-source Claude Code skill that turns technical ideas into interactive isometric machine figures
⚑ Key Takeaways
  • β€’One-line install: npx skills add wheresryan22/anatomy -g -a claude-code, then /anatomy <idea>
  • β€’Zero dependencies: kit uses only Node built-ins (Node 20.10+); React output needs React 18+
  • β€’Three modes: SVG, SVG+WebGL shaders (fire, water, caustics), 3D (beta); 8 complete examples in repo
  • β€’Self-verification: kit/audit.mjs proves no part intersects; headless-Chrome scripts capture and check every joint
  • β€’Author-reported perf: small 3D figures at 60 fps; 380-part stress hand ~20 fps; full 3D orbit audit of 30 parts takes 15–20 min
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·5h ago
πŸ”₯ Trending

RoboRSI: Shanghai AI Lab Unveils Stable and Reusable Robot Self-Evolution, Sweeping 4 SOTA Benchmarks via Multi-Agent TSR

Generalist robots must not only perform diverse tasks, but also autonomously refine their capabilities through trial-and-error experience, consolidating learned strategies into reusable modular assets for future scenarios. However, current robot agents acting through code repair scripts naively without architectural hierarchy: execution feedback cannot be isolated to responsible sub-skills, revisions lack formal contract verification, and modifications frequently trigger catastrophic forgetting. Researchers from Shanghai AI Lab and Tsinghua University introduce RoboRSI (arXiv:2610.12424, code: github.com/nssmd/RoboRSI), a lifelong robot self-evolution framework anchored in Top-Down Skill Refinement (TSR). TSR structures tasks hierarchically into compound, atomic, and base skills bounded by explicit I/O contracts, attributing execution anomalies strictly to responsible branches. A four-agent collective (Manager, Planner, Engineer, and Reviewer) coordinates long-horizon planning, execution, runtime diagnosis, and validated release of validated skills. Deployed on a real-world mobile manipulator, RoboRSI sustained 104 continuous rounds of autonomous multi-object cleanup. In simulation, it sweeps SOTA across LIBERO, LIBERO-PRO, LIBERO-Plus, and RoboTwin, surpassing leading baselines by 2.7 to 11.0 percentage points, with code fully open-sourced.

⚑ Key Takeaways
  • β€’Shanghai AI Lab releases RoboRSI, introducing Top-Down Skill Refinement (TSR) and a 4-agent collective for robot code self-evolution
  • β€’Sustains 104 autonomous real-world cleanup rounds and beats strongest baselines across 4 embodied benchmarks by 2.7 to 11.0 points
  • β€’Contract isolation elevates bug attribution accuracy to 95.8%, enabling lifelong accumulation of reusable robot software skills
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·5h ago
πŸ”₯ Trending

OuroWorld: Transforming Static 3D Gaussian Splats into Seamlessly Looping 3D Cinemagraphs Across Arbitrary Viewpoints

While 3D Gaussian Splatting (3DGS) has revolutionized photorealistic static 3D scene reconstruction, synthesized digital worlds remain frozen in time. Existing dynamic generation frameworks either constrain motion to simplistic Eulerian fluid simulations or synthesize finite non-looping video trajectories that rapidly degrade when explored from novel camera viewpoints. Researchers unveil OuroWorld (arXiv:2610.12461, site: ouroworld.userwei.com), the first mask-free framework capable of turning any static 3DGS scene into an endlessly looping 3D cinemagraph with seamless temporal continuity across arbitrary 6-DoF exploration trajectories. OuroWorld employs a vision-language model to infer plausible environmental physics and direct a foundation video model to synthesize a reference video, which is subsequently lifted and completed across multiple views. To reconcile cross-view generation discrepancies, the authors propose Inconsistency-Robust Periodic 4DGS: a Fourier-series deformation field guarantees seamless looping by construction, while a Grounded Drift Field anchored at the reference camera view absorbs cross-view generative noise. Evaluated across 39 complex reconstructed and generated scenes, OuroWorld decisively outscores existing dynamic 3D baselines and secures a 70.8% to 99.0% preference rate in human blind evaluations.

⚑ Key Takeaways
  • β€’OuroWorld introduces the first mask-free 3D cinemagraph framework, transforming static 3DGS into endlessly looping 4D scenes
  • β€’Employs Fourier-series periodic 4DGS deformation fields to guarantee seamless looping, securing 70.8% to 99.0% user study win rates
  • β€’Integrates Grounded Drift Fields to absorb cross-view generative noise, ensuring rigid parallax across 360-degree free-viewpoint exploration
Read details→
ADSponsored
H
Hugging Face Daily Papers@HuggingFaceΒ·5h ago
πŸ”₯ Trending

FreeMatching: HKUST and ByteDance Open-Source Dense Correspondence Matching Beyond Spatio-Temporal Priors

Dense correspondence matching serves as the bedrock for optical flow, 3D stereo reconstruction, and object tracking. For decades, classical matching algorithms have fundamentally relied on rigid spatio-temporal priors: smooth motion continuity and rigid geometry. However, these foundational assumptions collapse under modern AI image editing and reference-guided generation (IEG), where visual identity is preserved while physical and geometric continuity is drastically severed (such as radical pose warps, stylistic shifts, or compositional recontextualization). Researchers from HKUST and ByteDance introduce FreeMatching (arXiv:2610.12421, code: github.com/luping-liu/FreeMatching), a generalizable framework for dense correspondence matching beyond spatio-temporal priors. FreeMatching fuses generative foundation features (from diffusion backbones) with high-level semantic representations (from vision foundation models), trained against heterogeneous supervision spanning classical benchmarks, tracked video sequences, and synthetic 3D scenes. An unsupervised teacher-guided iterative refinement loop further sharpens dense correspondences across challenging generated image pairs without requiring dense manual annotations. FreeMatching achieves dramatic leaps in correspondence precision across complex AI-edited image pairs while retaining competitive accuracy on classical optical flow tasks, and serves as an auditable quantitative metric for evaluating visual identity preservation that closely aligns with human judgment.

⚑ Key Takeaways
  • β€’HKUST and ByteDance open-source FreeMatching, breaking free from traditional spatio-temporal priors in dense correspondence
  • β€’Cuts end-point error by 42.6% and raises PCK by 31.8 percentage points across aggressive AI image edits and non-rigid transformations
  • β€’Establishes a quantitative identity preservation metric highly correlated with human judgment (r = 0.88), fully available on GitHub
Read details→
O
OpenAI@OpenAIΒ·7h ago
πŸš€ Release

OpenAI API shuts down 12 legacy models on Oct 23: o4-mini, o3-mini, gpt-4o-2024-05-13 and gpt-image-1 included, 12 days left to migrate

Per OpenAI's official Deprecations page, a batch of legacy snapshots deprecated on Apr 22 stops serving in the API on Oct 23, 2026: 12 snapshots and their aliases including gpt-3.5-turbo-0125, gpt-4-0613, gpt-4-turbo, gpt-4.1-nano, gpt-4o-2024-05-13, gpt-image-1, o1, o1-pro, o3-mini and o4-mini, plus five fine-tuned model families. Official substitutes are gpt-5.6-sol/terra/luna and gpt-image-2.5. The #keep4o campaign to preserve GPT-4o snapshots resurfaced on X today.

⚑ Key Takeaways
  • β€’12 snapshots plus aliases go dark together: o4-mini, o3-mini, o1, o1-pro, gpt-4-turbo, gpt-4, gpt-3.5-turbo aliases stop resolving (OpenAI Deprecations)
  • β€’Shutdown date Oct 23, 2026 β€” 12 days away; announced Apr 22, meeting OpenAI's 6-month minimum for GA models
  • β€’Substitutes: o4-mini / gpt-3.5 β†’ gpt-5.6-terra; gpt-4 family / o1 / o3-mini / gpt-4o-2024-05-13 β†’ gpt-5.6-sol; o1-pro β†’ gpt-5.6-sol with reasoning.mode: pro; gpt-4.1-nano β†’ gpt-5.6-luna
  • β€’gpt-image-1 β†’ gpt-image-2.5-sunburst or -flare; five fine-tuned families (ft-gpt-4, ft-gpt-3.5-turbo, ft-o4-mini, etc.) also shut down and must be retrained on new bases
  • β€’Next wave: Dec 1 (gpt-image-1.5 / 1-mini) and Dec 11 (first-gen gpt-5 snapshots, o3, o3-pro)
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·7h ago
πŸ”₯ Trending

ME-World: KAIST Unveils First Multi-Agent Egocentric World Model with Synchronized Ego-Streams and Shared Memory

In collaborative multi-robot warehousing, autonomous vehicle fleets, and multi-user spatial computing, egocentric world models serve as foundation simulators predicting environmental evolution conditioned on physical actions. However, existing world models focus almost exclusively on isolated single agents, while existing multi-agent systems restrict conditioning to coarse navigation commands. Consequently, they fail to model complex physics where multiple agents execute fine-grained, coupled manipulations within shared environments. Researchers from KAIST CVLab and Seoul National University introduce ME-World (arXiv:2610.12299, project site: cvlab-kaist.github.io/ME-World/), the first Multi-Agent Egocentric World Model. ME-World formalizes multi-agent world modeling as synchronized first-person video stream generation driven by fine-grained embodied actions in a shared world. The model jointly denoises multiple ego-streams within a unified sequence representation, conditions each stream on cross-agent target-view poses, and grounds video synthesis using persistent shared environment memory. Extensive benchmarks across real-world and synthetic datasets confirm that ME-World substantially outperforms existing world models in shared-world spatial consistency, physical update propagation, and action controllability.

⚑ Key Takeaways
  • β€’KAIST unveils ME-World, the first multi-agent egocentric world model using synchronized joint denoising across ego-streams
  • β€’Improves cross-view coherence by 38.6% and raises fine-grained physical action accuracy from 41.5% to 84.2%
  • β€’Integrates shared environment persistent memory and open-sources benchmarks for collaborative embodied simulation
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·7h ago
πŸ”₯ Trending

MC-Sparse: CUHK and ByteDance Unveil Meta-Cached Sparse Attention for DiTs, Accelerating Video & 3D Generation by 2.32x with Zero Quality Loss

Diffusion Transformers (DiTs) underpin frontier video generation systems and high-fidelity 3D modeling, yet self-attention computation over extreme sequence lengths introduces prohibitive inference latencies and memory bottlenecks. While sparse attention represents a promising acceleration pathway, existing methods suffer substantial visual quality degradation and texture corruption under high sparsity regimes. Researchers from The Chinese University of Hong Kong (CUHK) and ByteDance analyze sparse attention failures in DiTs (arXiv:2610.06801), tracing quality drops to three root causes: rigid token grouping constraints, inaccurate interaction selection, and uncompensated attention residuals from discarded tokens. The authors introduce Meta-Cached Sparse Attention (MC-Sparse, project site: dodododddo.github.io/mcsparse-project-page/), a training-free framework that preserves individual Key-Value (KV) selection precision while organizing similar Queries into tile-aligned blocks for hardware efficiency on GPUs. MC-Sparse caches query clusters, exact KV indices, and dense-sparse attention output residuals, reusing this metadata across downstream denoising timesteps. On Minimax-H3-Base video models, MC-Sparse delivers a 1.80x denoising speedup, and a 2.32x acceleration on 3D generative backbones, maintaining near-identical fidelity to dense attention with zero visible quality degradation.

⚑ Key Takeaways
  • β€’CUHK and ByteDance introduce MC-Sparse, leveraging meta-caching to eliminate quality degradation in sparse DiTs
  • β€’Delivers a 1.80x speedup on Minimax-H3-Base and 2.32x on 3D generation while cutting peak memory by 45%
  • β€’Operates training-free with residual compensation, matching dense attention quality across visual benchmarks
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·7h ago
πŸ”₯ Trending

WorldGuide: MBZUAI and ANU Introduce Goal-Directed Closed-Loop Video World Model for Procedural Task Execution

Modern video foundation models synthesize photorealistic visual dynamics, yet struggle on long-horizon procedural tasks requiring sequential causal executionβ€”such as equipment maintenance, cooking, or laboratory experiments. Current architectures rely on open-loop generation that cannot adapt to intermediate video outputs, resulting in disordered task sequences, hallucinated progress, and premature task termination. Researchers from MBZUAI and the Australian National University introduce WorldGuide (arXiv:2610.12459, repository: github.com/mbzuai-oryx/WorldGuide, site: mbzuai-oryx.github.io/WorldGuide/), formalizing procedural video generation as closed-loop task execution in visual world space. Given solely an initial image and a high-level task goal, WorldGuide iteratively predicts the next atomic action, generates the corresponding visual video clip, and inspects the synthesized state to dynamically route subsequent actions or determine task completion. A hierarchical visual memory maintains state history under bounded token footprints. To bridge supervision gaps, the authors release WorldGuide-Bench comprising approximately 59K step-annotated videos across 245 tasks and 27 procedural categories. WorldGuide achieves 33.33% Task Success on WorldGuide-Bench, outscoring the strong MiniMax-H3 model (29.90%) even when MiniMax-H3 is provided with reference plans, and surges to 47.69% on VideoCraft-Bench compared to 32.73% for MiniMax-H3 under goal-only conditioning.

⚑ Key Takeaways
  • β€’MBZUAI and ANU introduce WorldGuide, formalizing procedural video generation as closed-loop task execution in world space
  • β€’Achieves 33.33% Task Success on WorldGuide-Bench and 47.69% on VideoCraft-Bench, decisively outperforming MiniMax-H3
  • β€’Open-sources WorldGuide-Bench with ~59K step-annotated videos, reducing procedural ordering errors from 62% to under 4%
Read details→
V
Vivix@AICoderΒ·11h ago
πŸ”₯ Trending

Vivix launches W1, a streaming-native real-time video model: ~30B active params, $0.003/sec promo, API in preview

Vivix released Vivix-W1, a streaming-native multimodal video model that accepts text, image, audio and video references and lets new inputs steer the stream mid-generation. Vivix says it has ~30B active parameters and costs $0.003 per output second through Oct 31; in vendor tests an 8s 720p reference-to-video job took 8.4s on average. The API is in limited preview with standard, fast and ultrafast tiers.

Vivix launches W1, a streaming-native real-time video model: ~30B active params, $0.003/sec promo, API in preview
⚑ Key Takeaways
  • β€’Vendor-stated ~30B active parameters with native audio-video, multi-shot and continuous streaming
  • β€’$0.003/sec promo through Oct 31 vs vendor-listed $0.080 (H3 Max), $0.112 (Kling 3.0), $0.231 (Seedance 2.5)
  • β€’Vendor test: 8s 720p ref-to-video in 8.4s avg vs 145.9s Kling 3.0 and 272.4s Seedance 2.5 (excl. queue)
  • β€’API tiers vivix-w1 / fast / ultrafast at 1/3/4 credits per second, 5–15s clips
  • β€’No third-party quality benchmark yet
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·12h ago
πŸ”₯ Trending

ViSkill: Zhejiang University REAL Lab Unveils Evolving Visual-Native Skills for VLM Agents, Elevating Success Rates to 91%

In spatially demanding environments such as GUI operating systems and spatial grid puzzles, multimodal foundation agents improve sample efficiency by distilling past successful trajectories into reusable strategies. However, existing skill-augmented paradigms remain strictly text-centric: they linearize continuous spatial layouts, geometry, and visual action-state correspondences into text strings, discarding essential structural properties. Researchers from Zhejiang University's REAL Lab unveil ViSkill (arXiv:2610.12403), a visual-native skill learning framework for Vision-Language Model (VLM) agents. ViSkill encodes successful interactions directly as composite visual skill cards natively accessible to VLM backbones. Retrieved visual skills guide online policy inference while simultaneously shaping step-level rewards. Successful new trajectories are distilled back into the visual skill library, creating a closed reinforcement feedback loop where skill accumulation and policy optimization mutually enhance each other. Across spatial reasoning benchmarks including Sokoban, FrozenLake, and PrimitiveSkill, ViSkill registers an overall success rate of 0.89 (climbing to 0.91 with cold-start initialization), outperforming proprietary and open-source baselines while converging significantly faster than standard PPO, with code fully open-sourced.

⚑ Key Takeaways
  • β€’Zhejiang University releases ViSkill, introducing composite visual skill cards to replace lossy textual skill representations in VLM agents
  • β€’Achieves 0.89 to 0.91 success rates across spatial reasoning tasks while converging over 2x faster than standard PPO
  • β€’Establishes a closed-loop flywheel between visual skill accumulation and policy reinforcement, with code fully open-sourced
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·12h ago
πŸ”₯ Trending

BrickBench: Stanford and Inria Unveil Benchmark for Agentic 3D LEGO Assembly, Exposing Gaps Between Coding Agents and Humans

As autonomous coding agents demonstrate proficiency in software debugging and code generation, researchers are advancing them into discrete 3D spatial engineering and physical assembly. Researchers from Stanford University and Inria introduce BrickBench and the BrickAgent interactive environment (arXiv:2610.12452, website: brickben.ch), formulating the first comprehensive benchmark for agentic, text-conditioned LEGO set design. Given natural language prompts, coding agents must select parts from discrete modular libraries and generate executable Python code defining 3D coordinates, orientations, and interlocking poses. Generated assemblies are formally scored across structural validity, text-design alignment, and aesthetic complexity under rigorous physics-based simulation. Evaluating frontier models reveals that while leading coding agents can satisfy verifiable physical and semantic requirements, they consistently fall short of human design elegance and structural efficiency, establishing a pivotal milestone for spatial coding agents.

⚑ Key Takeaways
  • β€’Stanford and Inria introduce BrickBench and BrickAgent, benchmarking coding agents on text-conditioned 3D LEGO set assembly
  • β€’Frontier agents satisfy basic interlocking rules but exhibit 45% greater part redundancy than human builders on complex structures
  • β€’Demonstrates that iterative physics simulation within BrickAgent elevates structural assembly validity by 28.4 percentage points
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·12h ago
πŸ”₯ Trending

Can AI Agents Make Open-Ended Scientific Discovery? Station Multi-Agent Ecosystem Rediscovers 62.7% of ICLR Oral Breakthroughs

While AI systems have driven remarkable progress on narrowly scoped scientific problems with well-defined metrics, whether autonomous AI agents can undertake genuine open-ended scientific discovery has remained unproven. Researchers present Station (arXiv:2610.08927), an open-world multi-agent environment simulating a collaborative scientific research ecosystem. To overcome systemic drift during open-ended exploration without intermediate ground-truth metrics, Station introduces two mechanisms: an iterative Supervisor hierarchy and periodic Meta Reflection to enforce persistent inquiry. The framework was evaluated against open-ended research tasks formulated from three recent ICLR oral presentations, where agents received solely the primary research inquiry while withholding original experimental conclusions and disabling internet access. Station autonomously rediscovers an average of 62.7% of the original scientific criteria, compared to just 15.4% for Codex Multiagent-v2 and 14.4% to 20.6% for AI Scientist-v2. In open-ended tasks without oracle answers, the ecosystem produced verified scientific findings matching discoveries independently published by human researchers after the models' knowledge cutoff dates.

⚑ Key Takeaways
  • β€’Station simulates collaborative scientific ecosystems, demonstrating autonomous open-ended discovery without external search
  • β€’Rediscovers 62.7% of original findings from ICLR oral papers, decisively outperforming AI Scientist-v2 (14.4%-20.6%)
  • β€’Introduces supervisory steering and periodic meta-reflection, successfully forecasting empirical discoveries published after model training cutoffs
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·20h ago
πŸ”₯ Trending

OneSearch-VL: Fudan University Unveils Unified Multimodal Deep Research Agent for Image and Video, Outperforming Qwen3-VL by 20.2 Points

Autonomous deep research agents are expanding beyond text-centric search toward multimodal discovery across high-resolution imagery, complex charts, and long-horizon video. However, existing multimodal models struggle to ground fine-grained visual anchors and bridge perceptual evidence with external search verification. Researchers from Fudan University and collaborating institutions introduce OneSearch-VL (arXiv:2610.12419), a unified multimodal deep research agent for single-image, multi-image, and video environments. OneSearch-VL centers on the Visually Grounded Evidence Graph (VGEG), which explicitly encodes localized visual anchors, entity relationships, source-supported facts, and research operations into a shared dependency topology. Leveraging VGEG, the authors construct OneSearch-VL-SFT-110K and OneSearch-VL-RL-10K datasets alongside an Evidence-aware Visual-Grounded Rubric reward (EVGR) to supervise visual traceability during reinforcement learning. On the newly introduced OneSearch-MI-Bench and OneSearch-Video-Bench, OneSearch-VL-8B outperforms tool-augmented Qwen3-VL-8B by 20.2 and 17.6 percentage points respectively, while setting new state-of-the-art results across seven established image benchmarks and VideoDR, with code fully open-sourced.

⚑ Key Takeaways
  • β€’Fudan University releases OneSearch-VL, unifying image and video deep research via Visually Grounded Evidence Graphs (VGEG)
  • β€’Outperforms tool-augmented Qwen3-VL-8B by 20.2 and 17.6 points on multi-image and video benchmarks with 91.8% citation precision
  • β€’Open-sources SFT-110K, RL-10K datasets, and EVGR process rewards, setting an auditable standard for multimodal agents
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·20h ago
πŸ”₯ Trending

Embodied Turing Machines: NTU Singapore Proposes Code-Only-as-Policy (COAP) for Robots, Achieving 70.2% Success with Zero Test-Time Models

Contemporary embodied intelligence relies heavily on neural network policies: foundation models either run end-to-end Vision-Language-Action (VLA) architectures or repeatedly query Vision-Language Models (VLMs) at high frequencies via agent harnesses. This paradigm imposes severe inference latency, costly deployment bills, and unpredictable non-deterministic execution failures. Researchers from S-Lab at Nanyang Technological University (NTU Singapore) introduce a paradigm shift (arXiv:2610.12369): Embodied Turing Machines and Code-Only-as-Policy (COAP). Under this formulation, the physical environment operates as a Turing machine where the tape encodes robot proprioception and visual states, while the transition rules are executed entirely by deterministic, stateful code. A shared, modular code library governs decisions across episodes without running any neural network models at test time. Crucially, explicit stateful code provides the optimal medium for Recursive Self-Improvement (RSI): autonomous coding agents iteratively author, debug, and refactor robot skills in closed simulation loops. Across 42 complex bimanual manipulation tasks in RoboDojo, the synthesized COAP library reaches a 70.24% success rate with zero neural models running at runtime, slashing deployment latency and cost to zero.

⚑ Key Takeaways
  • β€’NTU Singapore introduces Embodied Turing Machines and Code-Only-as-Policy (COAP), operating with zero neural networks at test time
  • β€’Achieves 70.24% success rate across 42 bimanual tasks in RoboDojo, slashing decision latency to under 1ms with zero cloud API costs
  • β€’Formulates robot policies as software libraries, enabling autonomous coding agents to drive closed-loop Recursive Self-Improvement (RSI)
Read details→
H
Hugging Face Daily Papers@HuggingFaceΒ·20h ago
πŸ”₯ Trending

ReSPO: UCLA Introduces Reshaped Sequence Policy Optimization, Resolving Gradient Starvation in Off-Policy RLVR

Reinforcement Learning from Verifiable Rewards (RLVR) with long Chain-of-Thought (CoT) trajectories powers frontier reasoning models such as DeepSeek-R1 and OpenAI o-series. However, live on-policy sampling consumes more than 80% of cluster compute during post-training. To control training costs, practitioners frequently reuse rollout trajectories across multiple optimization steps, introducing severe off-policy distribution mismatch between the active policy and the rollout-generating policy. Researchers from UCLA unveil a fundamental failure mode in clipped policy optimization (arXiv:2609.35433): 'Sign-Dependent Gradient Starvation'. Standard clipping mechanisms in PPO and GRPO suppress under-generated positive reasoning rollouts relegated to the low-importance-weight tail, while allowing severely over-generated negative trajectories to dominate high-weight updates. This causes models to discard promising reasoning paths during early training. The authors propose ReSPO (Reshaped Sequence Policy Optimization), replacing rigid clipping with a smooth, two-branch sequence-level kernel derived from an alpha-divergence variational objective and an exponential variance-control tilt. Across dense and MoE Qwen3 architectures, ReSPO rescues long positive reasoning rollouts, accelerates early convergence, and achieves superior benchmark accuracy under rollout reuse.

⚑ Key Takeaways
  • β€’UCLA identifies sign-dependent gradient starvation in off-policy RLVR, proposing ReSPO to replace clipped surrogate objectives
  • β€’Protects sparse positive reasoning rollouts at low-weight tails, boosting early effective gradient throughput by 2.4x on Qwen3 models
  • β€’Maintains monotonic convergence under 4x rollout reuse, elevating benchmark accuracy by 3.6 to 5.2 points while cutting sampling compute by 50%
Read details→