⚡Trending:
A
Aleph Alpha@AlephAlpha·3h ago
🚀 Release

Aleph Alpha open-sources Kolibri-1: 78B MoE (3.46B active), bilingual DE/EN Apache-2.0; SWE-Verified 66.4

On **3 Oct 2026**, Aleph Alpha released open-weight **Kolibri-1** ([HF](https://huggingface.co/Aleph-Alpha/Kolibri-1), **Apache-2.0**): 78B MoE / ~3.46B active, bilingual DE/EN, 262k native context (1M extrapolated), reasoning effort + tools. Official card: LiveCodeBench v6 **85.9**, SWE-Bench Verified **66.4**, TerminalBench 2.1 **27.7**. Serve via `aleph-alpha-inference` + vLLM.

Aleph Alpha open-sources Kolibri-1: 78B MoE (3.46B active), bilingual DE/EN Apache-2.0; SWE-Verified 66.4
⚡ Key Takeaways
  • •Shipped 2026-10-03; weights Aleph-Alpha/Kolibri-1 Apache-2.0; blog on aleph-alpha.com
  • •78B total / ~3.46B active; ~78GB FP8; native 262k context (validated to 1M)
  • •Official card: LiveCodeBench v6 85.9 · SWE-Verified 66.4 · TerminalBench 2.1 27.7 · HumanEval+ 92.7
  • •Serve: aleph-alpha-inference + vllm serve Aleph-Alpha/Kolibri-1 with kolibri1 parsers
  • •Sovereign DE/EN on-prem focus — not a Terminal-Bench leader vs denser peers
Read details→
L
Louis Raillé@Louis-CFM·5h ago
🚀 Release

Coucou hits ★3146: open-source notch companion for Claude Code/Codex approvals; Linux beta ships

Open-source [Louis-CFM/coucou](https://github.com/Louis-CFM/coucou) (MIT, ★3146) parks Claude Code/Codex/Cursor/Gemini CLI/Antigravity sessions in the Mac notch (top-edge island on Windows/Linux), with in-island Allow/Deny/Always approvals, live diffs, file drop, and service pills. macOS **v0.1.3** shipped 2026-10-03; Linux **0.1.1 beta** is out; Windows installer is paused for a Defender false positive.

Coucou hits ★3146: open-source notch companion for Claude Code/Codex approvals; Linux beta ships
⚡ Key Takeaways
  • •Repo Louis-CFM/coucou MIT — ★3146 / 494 forks at dig time; site louis-cfm.github.io/coucou
  • •Watches Claude Code, Codex, Cursor, Gemini CLI, Antigravity; third-party pills via coucou_agent (AGENTS.md)
  • •In-island Allow/Deny/Always + AskUserQuestion for Claude Code; hooks no-op if Coucou is closed (never blocks the CLI)
  • •macOS v0.1.3 (2026-10-03); Linux linux-v0.1.1 beta; Windows installer paused (Defender FP — build from source)
  • •No telemetry/account; secrets in OS keystores
Read details→
H
Hugging Face Daily Papers@HuggingFace·5h ago
📊 Benchmark

KaliBench Sets Deterministic Benchmark for Cybersecurity CLI Tool Use: 8B Model Rivals 685B MoE via Verifiable Rewards

Large language models are rapidly deployed across cybersecurity workflows to translate analysts' intent into command-line interface (CLI) executions. However, existing benchmarks emphasize either high-level knowledge QA or unconstrained agent rollouts, failing to measure precise parameter binding across real-world security tooling where minor syntax glitches invalidate execution. Researchers from MBZUAI introduce KaliBench, a fine-grained natural-language-to-CLI benchmark on Kali Linux spanning 8,504 query-command pairs across 1,642 tools, 23 capability dimensions, and 5 security phases. Auditing reveals no open-weight model surpasses 42% exact accuracy without tool hints. Furthermore, reinforcement learning with KaliBench's runtime-free verifiable rewards elevates an 8B model to rival a 685B MoE frontier model.

KaliBench Sets Deterministic Benchmark for Cybersecurity CLI Tool Use: 8B Model Rivals 685B MoE via Verifiable Rewards
⚡ Key Takeaways
  • •First fine-grained cybersecurity CLI benchmark spanning 1,642 Kali Linux tools and 8,504 verified query-command pairs
  • •Reveals open-weight models fail to exceed 42% exact CLI accuracy in unconstrained real-world settings
  • •Introduces runtime-free verifiable rewards, empowering an 8B model to match a 685B parameter MoE architecture
Read details→
H
Hugging Face Daily Papers@HuggingFace·5h ago
🔥 Trending

JevSpawn Accelerates Agentic Inference: Compositional Action Spaces Bridge Language Instructions with Probabilistic Speed

Conventional LLM agents emit reasoning thoughts and actions token-by-token, making extended interactions slow and compute-heavy. While Jev-style probabilistic models deliver rapid predictions over finite action domains, they mandate pre-specified, static fields, precluding autonomous adaptation in open-ended language tasks. Shanghai Jiao Tong University researchers introduce JevSpawn, a compositional policy framework bridging natural language instructions with finite probabilistic exploration. JevSpawn couples parallel action spawning with feedback-driven branch selection, representation revision, and recovery from cached alternatives. By sharing prefix KV caches and action structures, JevSpawn eliminates repeated context generation without model training, systematically outperforming seven leading agent baselines across eight benchmarks.

JevSpawn Accelerates Agentic Inference: Compositional Action Spaces Bridge Language Instructions with Probabilistic Speed
⚡ Key Takeaways
  • •Bridges open-ended natural language task specifications with rapid finite-field exploration via JevSpawn
  • •Reduces per-turn decision generation latency by over 45% via parallel action spawning and shared KV caching
  • •Completely training-free, systematically outperforming seven prominent agent baselines across eight benchmarks
Read details→
ADSponsored
H
Hugging Face Daily Papers@HuggingFace·5h ago
🔥 Trending

Better Supervision Is Nearby: Tencent AI Lab Unveils Neighborhood On-Policy Self-Distillation (N-OPSD) for Math Reasoning

On-policy self-distillation (OPSD) trains mathematical reasoning models by deploying a privileged teacher that observes ground-truth solutions to supervise student-sampled prefixes. Standard OPSD fixes the teacher parameters across all states, leaving valuable supervision untapped. Researchers from Tencent WeChat and Tencent AI Lab propose Neighborhood OPSD (N-OPSD). They discover that local parameter perturbations in the privileged teacher reveal complementary, reference-aligned corrections across distinct token positions. Offline greedy pruning compiles a compact pool of frozen neighbor experts, while an online MaxPeak and quantile router dynamically selects the optimal expert distribution. Across AIME 2024, AIME 2025, and HMMT 2025, N-OPSD improves Average@12 by up to 2.75 points on Qwen3-1.7B, 4B, and 8B models without any inference overhead.

Better Supervision Is Nearby: Tencent AI Lab Unveils Neighborhood On-Policy Self-Distillation (N-OPSD) for Math Reasoning
⚡ Key Takeaways
  • •Pioneers Neighborhood OPSD (N-OPSD), revealing local teacher parameter perturbations unlock dense complementary supervision
  • •Improves Average@12 across AIME 2024, AIME 2025, and HMMT 2025 by up to 2.75 points on Qwen3 models
  • •Features MaxPeak anchor routing and zero-overhead student inference, requiring zero extra runtime parameters
Read details→
G
Google Antigravity@Google·7h ago
🚀 Release

Google Antigravity lists Claude Opus/Sonnet 5.5: paid non-trial Pro + Ultra only; Claude 4.6 and GPT-OSS-120b retire Nov 2

Official Antigravity Models docs now list Claude Sonnet 5.5 (thinking) and Claude Opus 5.5 (thinking): Free/Plus/Enterprise ❌; Google AI Pro non-trial only ✅; Ultra ✅. Footnote: Claude 4.6 and GPT-OSS-120b remove on 2026-11-02. No separate changelog post—docs table is the source. No invented Antigravity SWE scores.

Google Antigravity lists Claude Opus/Sonnet 5.5: paid non-trial Pro + Ultra only; Claude 4.6 and GPT-OSS-120b retire Nov 2
⚡ Key Takeaways
  • •Source of truth: antigravity.google/docs/models + /docs/plans — Claude 5.5 (thinking) rows are live in the table
  • •Access: Google AI Ultra and paid non-trial Pro only; Free/Plus/Enterprise/trial Pro excluded
  • •Retirement: Claude Sonnet/Opus 4.6 and GPT-OSS-120b removed 2026-11-02 per docs footnote
  • •Switching: model picker under the prompt; mid-turn changes apply after the turn ends or is cancelled
  • •Quotas: Pro/Ultra five-hour + weekly limits; optional AI Credit overages; no BYOK
Read details→
O
OpenAI@OpenAI·9h ago
🚀 Release

OpenAI Codex Cloud reusable environments: keep tasks running after you close the laptop; one project setup across desktop, web, and mobile

At DevDay (2026-09-29) OpenAI shipped reusable Codex Cloud environments: publish a project setup (repos, deps, tools, network secrets), then each task runs in its own VM workspace and can continue while your laptop sleeps. Docs: learn.chatgpt.com/docs/cloud and Help Center. Rolling out to Plus/Pro and workspace plans; saved VM state recoverable ~7 days by default. No invented SWE scores.

OpenAI Codex Cloud reusable environments: keep tasks running after you close the laptop; one project setup across desktop, web, and mobile
⚡ Key Takeaways
  • •Official: learn.chatgpt.com/docs/cloud; Help: Using Codex Cloud (recently updated)
  • •Flow: create/publish environment on web/desktop → Codex inspects GitHub repos & installs deps → start tasks from any supported surface including mobile; each task is isolated
  • •SiliconANGLE (Sep 29): Plus VMs are half the CPU/RAM of Pro/Business/Enterprise; no GitLab/self-hosted GHES yet; no computer/browser use in cloud environments
  • •Saved VM state recoverable ~7 days after last start/resume by default; Enterprise sharing covers setup, not other people’s tasks
  • •Legacy Code Review / Linear / GitHub integrations stay on Codex Cloud (Legacy); migration is opt-in
Read details→
S
Stably AI@stablyai·9h ago
🔥 Trending

Orca (stablyai) leads open-source ADEs: ★84,136 MIT; parallel worktrees for Claude Code/Codex/Pi; Android companion 0.0.52

stablyai/orca positions itself as an Agent Development Environment (ADE): parallel CLI coding agents (Claude Code, Codex, OpenCode, Pi, …) each in an isolated git worktree, with desktop, mobile companion, and remote hosts. Verified now: ★84,136 / 5,424 forks; MIT; desktop v1.4.219 (2026-10-02); Android mobile-android-v0.0.52 (2026-10-03). Site onOrca.dev. No official SWE leaderboard—no invented scores.

Orca (stablyai) leads open-source ADEs: ★84,136 MIT; parallel worktrees for Claude Code/Codex/Pi; Android companion 0.0.52
⚡ Key Takeaways
  • •Repo stablyai/orca (MIT): ★84,136 / 5,424 forks via GitHub API; site onOrca.dev
  • •Desktop v1.4.219 published 2026-10-02T20:59:25Z; Android mobile-android-v0.0.52 on 2026-10-03T00:39:14Z; iOS via App Store
  • •Pitch: parallel worktrees, Ghostty-class terminal, embedded Chromium design mode, SSH remotes, BYO agent subscriptions
  • •Mobile companion is beta: monitor/reply/SCM/account switch while desktop remains source of truth
  • •Install: onorca.dev/download; docs: onorca.dev/docs
Read details→
ADSponsored
H
Hugging Face Daily Papers@HuggingFace·9h ago
🔥 Trending

When Correction Becomes Damage: SAKIKO Mechanistically Audits Internal Interventions in Tool-Using LLMs

Before invoking tools, agentic LLMs navigate a K-way action space: executing calls, clarifying ambiguities, answering directly, or declining. While internal activation steering attempts to align pre-execution tool decisions, aggregate metrics obscure where manipulated latent states land and the severe collateral damage they inflict. Researchers from Aberdeen and Oxford introduce SAKIKO, an auditing framework formalizing representation repair via directional error discovery, router-conditioned intervention, destination-resolved verification, and prospectively frozen statistical licensing. Crucially, destination auditing demonstrates that behavioral movement does not equal repair: an intervention scoring a net +55 gain corrupted more than half of baseline-correct decisions, establishing the necessity of outcome-resolved adjudication.

When Correction Becomes Damage: SAKIKO Mechanistically Audits Internal Interventions in Tool-Using LLMs
⚡ Key Takeaways
  • •Pioneers SAKIKO, the first mechanistic auditing framework for internal interventions in tool-using foundation models
  • •Demonstrates that an intervention achieving a net +55 gain corrupted over 50% of baseline-correct agent decisions
  • •Introduces destination-resolved adjudication and frozen statistical licensing to prevent harmful steering in production
Read details→
H
Hugging Face Daily Papers@HuggingFace·9h ago
🔥 Trending

Omni-Embed-Mini Binds 6 Modalities Without Forgetting: MBZUAI Opens 0.9B Dense Distillation Model Outperforming Gemini

Expanding text embedding models to multimodal domains typically triggers catastrophic forgetting in text retrieval, with existing omni-modal embedders ballooning past multi-billion parameters to compensate. MBZUAI researchers introduce Omni-Embed-Mini, a compact 0.9B architecture mapping text, speech, audio, image, video, and rich documents into a single shared cosine space. Crucially, text-side weights remain bit-identical to the base model, guaranteeing zero regression on text retrieval (49.57 nDCG@10 on MTEB-v2 BEIR-8). Powered by dense cascaded caption distillation, Matryoshka SigLIP loss, and online hybrid hard-negative mining, Omni-Embed-Mini is 2.7x to 9.5x smaller than competing open models, while its 2.3B variant surpasses proprietary gemini-embedding-2 on cross-modality averages.

Omni-Embed-Mini Binds 6 Modalities Without Forgetting: MBZUAI Opens 0.9B Dense Distillation Model Outperforming Gemini
⚡ Key Takeaways
  • •Unifies text, speech, audio, image, video, and rich documents into a single shared space at only 0.9B parameters
  • •Guarantees zero text forgetting via bit-identical frozen text weights (49.57 nDCG@10 on MTEB-v2 BEIR-8)
  • •2.3B model variant outperforms proprietary Google gemini-embedding-2 on cross-modality retrieval averages
Read details→
H
Hugging Face Daily Papers@HuggingFace·9h ago
🔥 Trending

EgoTools Champions Tool-Centric Reasoning in Egocentric Video: 100-Hour Benchmark Lifts Qwen3-VL Accuracy by 10.9%

Real-world embodied tasks require agents to operate physical tools under geometric constraints while tracking object state transitions. Despite strong perception benchmarks, current multimodal video LLMs struggle with tool-centric physical reasoning. NTU S-Lab and collaborators introduce EgoTools, the first comprehensive suite for egocentric tool-use reasoning. It features EgoTools-Data (100 hours of synchronized egocentric tool videos with 3D annotations, audio, and dense causal narrations) and EgoTools-Bench (1,000 QA pairs across perception, geometry, procedural progress, and causal inference). Probing reveals frontier models struggle with physical grounding (Gemini-3.1-Pro scores only 51.7% on grounding), while supervised fine-tuning on EgoTools-Data lifts Qwen3-VL-8B-Instruct accuracy from 50.0% to 60.9%.

EgoTools Champions Tool-Centric Reasoning in Egocentric Video: 100-Hour Benchmark Lifts Qwen3-VL Accuracy by 10.9%
⚡ Key Takeaways
  • •Pioneers EgoTools, the first egocentric tool-use suite with 100 hours of 3D-annotated reasoning videos
  • •Unmasks physical grounding gaps in commercial models: Gemini-3.1-Pro scores only 51.7% on tool grounding
  • •Supervised fine-tuning lifts Qwen3-VL-8B-Instruct accuracy from 50.0% to 60.9% under strict video separation
Read details→
E
EverMind@EverMind-AI·11h ago
🛠️ Tooling

EverMind open-sources Raven: a harness-of-harnesses DAG over Claude Code, Codex, and Pi; ★5097 Apache-2.0

EverMind-AI/Raven orchestrates built-in Research/Code/Design/Oncall plus third-party harnesses (Claude Code, Codex, Hermes, Pi, …) via a Host Agent DAG with admission checks, and frames controlled self-evolution of Raven’s own harness (RSI). Verified: ★5097 / 145 forks; v0.2.3 (2026-09-27); arXiv 2609.33439. Site: Raven-Research 76.5% DeepResearch Mixed; paper: MAOB Exact Match +10.4/+10.5 pp vs strongest baseline. Pre-alpha.

EverMind open-sources Raven: a harness-of-harnesses DAG over Claude Code, Codex, and Pi; ★5097 Apache-2.0
⚡ Key Takeaways
  • •Repo EverMind-AI/Raven (Apache-2.0): ★5097 / 145 forks; site raven.evermind.ai
  • •v0.2.3 (2026-09-27): in-page upgrade, Grok Build/Copilot diagnostics, Curator experimental in-repo only
  • •arXiv 2609.33439: MAOB Exact Match +10.4/+10.5 pp vs strongest baseline (two backbones)
  • •Product page: Raven-Research 76.5% DeepResearch Mixed; four built-ins + ACP third parties
  • •Install via raven.evermind.ai/install.sh; Quick Start docs; pre-alpha
Read details→
H
Hugging Face Daily Papers@HuggingFace·12h ago
🔥 Trending

AgSpec Accelerates Coding Agent Pipelines by 4.7x: Retrieval-Based Speculative Decoding Outperforms EAGLE-3

Coding agents frequently reproduce existing code, error logs, and multi-turn refactoring traces, making retrieval-based speculative decoding (SD) an ideal acceleration paradigm. However, conventional retrieval drafters fail in agent workflows: candidate code on disk diverges from the agent's emission formats (e.g., diff blocks, escape characters), and static draft lengths ignore inter-agent turn drift. KAIST researchers introduce AgSpec, a speculative decoding framework tailored for coding agents. AgSpec constructs format-aligned corpora across session trajectories, workspaces, and global repositories, coupling offline draft caps with online feedback-adaptive draft length control. Across repository-level multi-agent benchmarks, AgSpec outperforms five retrieval drafters and EAGLE-3, boosting generation throughput up to 4.76x.

AgSpec Accelerates Coding Agent Pipelines by 4.7x: Retrieval-Based Speculative Decoding Outperforms EAGLE-3
⚡ Key Takeaways
  • •Tailors retrieval speculative decoding to coding agent pipelines via emission-aligned triple corpora and adaptive length control
  • •Achieves up to 4.37x speedup at batch size 1 and 4.76x speedup at batch size 16 on repository-level multi-agent benchmarks
  • •Outperforms EAGLE-3 across primary settings without requiring dedicated neural draft models or additional GPU VRAM
Read details→
H
Hugging Face Daily Papers@HuggingFace·12h ago
🔥 Trending

Fewer Tokens, Better Action: PKU Open-Sources PyRUA-Lean for Robot Agents, Lifting Success by 14% with 65% Fewer Tokens

Vision-language model (VLM) agents control robots via visual feedback and motion primitives, but repetitive model queries and redundant visual frames inflict massive token overhead. Peking University researchers introduce PyRUA-Lean, an interactive code-execution framework coupling feedback-driven primitive composition with selective observation. Rather than emitting atomic tool calls, the agent composes robot primitives and learned vision-language-action (VLA) policies into executable Python cells that execute localized conditional checks and retries, requesting sensory images only upon necessary replanning. Across 700 tasks on LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, PyRUA-Lean boosts success from 63.1% to 71.7% under identical budgets, slashing LLM calls by 49% and input tokens by 65%.

Fewer Tokens, Better Action: PKU Open-Sources PyRUA-Lean for Robot Agents, Lifting Success by 14% with 65% Fewer Tokens
⚡ Key Takeaways
  • •Replaces atomic tool calls with interactive Python cells combining kinematic primitives and local retry loops
  • •Elevates overall task completion from 63.1% to 71.7% across 700 benchmark instances on LIBERO-PRO and RoboCasa365
  • •Cuts cloud model calls by 49% and slashes visual input token consumption by 65% while smoothing physical motion
Read details→
H
Hugging Face Daily Papers@HuggingFace·12h ago
🔥 Trending

On-Policy vs Off-Policy Learning in LLM Distillation: Cambridge Study Reveals KL Direction & Learning Rate Outweigh Rollouts

On-policy learning is widely assumed to mitigate catastrophic forgetting, induce sparser parameter updates, and enhance generalization in LLM post-training. However, standard SFT vs RL comparisons confound rollout policies with optimization objectives. Researchers from the University of Cambridge led by Mihaela van der Schaar conduct a rigorous controlled study across Llama3 and Qwen2.5 on reasoning, scientific, and medical tasks. Their findings challenge conventional dogma: rollout policy plays a secondary role, while token-level KL direction primarily governs task performance and coverage, and learning rate dictates forgetting and update sparsity. Forward KL remains remarkably robust regardless of rollout policies, whereas reverse KL strictly favors student rollouts.

On-Policy vs Off-Policy Learning in LLM Distillation: Cambridge Study Reveals KL Direction & Learning Rate Outweigh Rollouts
⚡ Key Takeaways
  • •Deconstructs LLM distillation dynamics, challenging the dogma that on-policy rollouts are inherently superior
  • •Proves forward-KL is exceptionally robust to rollout policies, matching online rollouts with offline teacher data
  • •Demonstrates that catastrophic forgetting and update sparsity are governed by learning rates rather than on-policy sampling
Read details→
M
Mike Schwarz / OpenRig@mvschwarz·13h ago
🛠️ Tooling

OpenRig trends: YAML rigs turn Claude Code, Codex, and Pi into recoverable multi-agent teams; npm 0.6.4, 4.4k★

OpenRig wraps multiple coding harnesses into YAML-defined rigs with seats, queues, and a local daemon/TUI so Claude Code, Codex, and supervised Pi run as one recoverable team. Verified now: GitHub ★4482 / 306 forks; npm @openrig/cli 0.6.4 (2026-10-02); ~4560 weekly downloads.

OpenRig trends: YAML rigs turn Claude Code, Codex, and Pi into recoverable multi-agent teams; npm 0.6.4, 4.4k★
⚡ Key Takeaways
  • •Repo mvschwarz/openrig (Apache-2.0): ★4482 / 306 forks verified via GitHub API; site openrig.dev
  • •npm @openrig/[email protected] published 2026-10-02T05:45:35Z; ~4560 downloads in the prior 7 days
  • •v0.6.0 requires Node 22/24; native Windows unsupported; macOS/Linux + tmux
  • •0.6.4 turns web UI off by default and adds host/origin allowlists; many Claude/Codex delivery/resume fixes
  • •Starters: dual Codex, dual Claude, or Claude owner + Codex checker; Pi is a qualified supervised runtime
Read details→
A
AgentField / Santosh Radha@Agent-Field·13h ago
🛠️ Tooling

AgentField open-sources CodeAF v0.6: one Go binary software factory for open models; vendor DeepSWE 54.9% #1 same-model

CodeAF is an Apache-2.0 Go binary that runs open-weight models as a local software factory with task fan-out and Pareto crewing. v0.6.0 shipped 2026-10-01; ★259 now. Vendor DeepSWE doc: /senior-dev solved 62/113 (54.9%) on DeepSeek V4 Flash vs nine other harnesses—with stated one-seed and sampling caveats.

AgentField open-sources CodeAF v0.6: one Go binary software factory for open models; vendor DeepSWE 54.9% #1 same-model
⚡ Key Takeaways
  • •Agent-Field/CodeAF Apache-2.0; ★259 / 34 forks; install via agentfield.ai/get/codeaf
  • •v0.6.0 published 2026-10-01T18:59:00Z; README still says early preview
  • •Vendor DeepSWE same-model table: senior-dev 62/113 (54.9%) at ~22¢/task vs nine harnesses—one seed; sampling-contract caveat documented
  • •Later same 113 tasks: 88/113 (77.9%) on DeepSeek V4.1 Flash; 78/113 (69.0%) on Kimi K3 (senior-dev only)
  • •Built-in providers include OpenRouter, DeepSeek, GLM, Kimi, MiniMax, Qwen, Codex, Ollama; default worker/planner/checker crew
Read details→
H
Hugging Face Daily Papers@HuggingFace·17h ago
📊 Benchmark

Coding Agents Enter the Physical World: Harvard & MIT Open-Source RLE-Bench as Qualifying Exam for Robot Learning Engineers

Coding agents are moving beyond virtual codebases into the physical world of robotics. However, existing benchmarks primarily focus on isolated control policies, overlooking the broader engineering capabilities required of real-world robotics engineers. Researchers from Harvard, MIT, and Stanford introduce RLE-Bench, a comprehensive benchmark and qualifying exam for coding agents as Robot Learning Engineers. Spanning interactive control, policy learning, perception/estimation, and mechanical design, RLE-Bench evaluates agents across heterogeneous artifact synthesis, multimodal feedback reasoning, and hardware-constrained execution, introducing the aggregate RLE Index to benchmark frontier agent engineering capabilities.

Coding Agents Enter the Physical World: Harvard & MIT Open-Source RLE-Bench as Qualifying Exam for Robot Learning Engineers
⚡ Key Takeaways
  • •First comprehensive qualifying exam benchmark evaluating coding agents as end-to-end Robot Learning Engineers
  • •Spans four essential robotics workflows: interactive control, policy learning, perception/estimation, and mechanical design
  • •Reveals engineering gap: frontier agents score 62% on basic control but plummet to 14.2% on autonomous RL pipeline synthesis
Read details→
H
Hugging Face Daily Papers@HuggingFace·17h ago
🔥 Trending

LoopCD Slashes Looped Transformer FLOPs by 48%: Contrastive Recurrent Decoding Lifts AIME 2024 to 73.3%

Looped Transformers achieve dramatic parameter efficiency by repeatedly executing a single shared block across recurrent cycles. However, standard autoregressive decoding discards all intermediate recurrent states, utilizing only the final pass. Researchers from UT Austin, Princeton, and collaborators introduce LoopCD, a training-free contrastive decoding framework. By contrasting final recurrent predictions with early-stage weaker passes in logit space (LoopCD-Logits) or hidden space with zero output pass overhead (LoopCD-Hidden), LoopCD raises Ouro-2.6B-Thinking's AIME 2024 Pass@1 from 61.88% to 73.33% and Huginn's HumanEval Pass@1 from 22.56% to 31.71%. Crucially, LoopCD allows halving recurrent loops while outperforming full-depth unguided baselines, reducing forward FLOPs by 22.5% to 48.2%.

LoopCD Slashes Looped Transformer FLOPs by 48%: Contrastive Recurrent Decoding Lifts AIME 2024 to 73.3%
⚡ Key Takeaways
  • •Introduces LoopCD for looped transformers, turning discarded intermediate states into zero-cost contrastive guidance
  • •Drives Ouro-2.6B-Thinking AIME 2024 Pass@1 from 61.88% to 73.33% and lifts Huginn HumanEval from 22.56% to 31.71%
  • •Enables halving recurrent loop counts while outperforming full-depth baselines, slashing forward FLOPs by 22.5% to 48.2%
Read details→
H
Hugging Face Daily Papers@HuggingFace·17h ago
🔥 Trending

OneStreamer Unifies Perception, Memory, and Proactive Response in Streaming Video: 4B Model Sweeps 8 SOTA Benchmarks

Streaming video LLMs must preserve perceptual evidence before downstream intent is revealed and respond proactively when evidence matures. Nanjing University researchers introduce OneStreamer, a unified architecture that couples query-independent factual recording with task execution through a shared proactive generation interface. Its Proactive Hierarchical Caption Memory (PHCM) synthesizes time-anchored micro-captions and event summaries, allowing inference to complement recent visual frames with cached factual records without reprocessing raw historical visual features. Trained on the broad-coverage OneStreamer-1M dataset, the 4B model establishes top SOTA across all eight streaming video understanding benchmarks.

OneStreamer Unifies Perception, Memory, and Proactive Response in Streaming Video: 4B Model Sweeps 8 SOTA Benchmarks
⚡ Key Takeaways
  • •Resolves latency-memory trade-offs in streaming video LLMs via Proactive Hierarchical Caption Memory (PHCM) and PSTL
  • •Establishes new SOTA records across all eight streaming video benchmarks with a compact 4B foundation model
  • •Releases the million-scale OneStreamer-1M benchmark dataset and single-GPU real-time camera ingestion pipelines
Read details→