arXiv 2610.02163 (PDF stamp 2026-10-02) and autocompact.github.io introduce AutoCompact: a proactive compact() action in a terminal REPL scaffold, trained via judge-corrected on-policy SFT then outcome GRPO so the policy learns when to compact, what to keep, and how to continue. Base Qwen3-Coder-30B-A3B-Instruct; SWE-bench Verified 39.6% vs 30.4% base (+9.2 pp); SWE-PolyBench Verified 24.5% vs 19.5% (+5.0 pp). No public weights/code repo found.

Key Takeaways

  • ✓Sources: arXiv 2610.02163 + autocompact.github.io; PDF Last-Modified 2026-10-02
  • ✓Mechanism: proactive compact() → # Auto Context Summary; task-phase vs length-triggered
  • ✓Training: 1052 judge-corrected SWE-rebench trajectories for SFT; then SWE-Gym GRPO with binary patch-pass reward only
  • ✓Results (3-run avg, same scaffold): SWE-bench Verified 39.6% vs 30.4% base; PolyBench 24.5% vs 19.5%; after RL compact() on 58.5% of tasks
  • ✓Caveat: no public weights or standalone code repo found — method paper, not a drop-in product
🧭

Heavy Claude Code use: compare subscription limits and API bills

Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Background

Repo-level coding agents accumulate stale exploration—failed hypotheses, long tool dumps, ruled-out paths. Length-triggered compaction often fires mid-stage. AutoCompact (arXiv 2610.02163; autocompact.github.io) treats compaction as a learned policy action via proactive compact().

Mechanism

Terminal REPL plus compact() → # Auto Context Summary, keeping the original task and recent turns. GPT-5.5-Codex judge corrects trigger/summary/continuation online on SWE-rebench (1052 SFT trajectories), then SWE-Gym GRPO with binary patch-pass reward only—no compaction-specific shaping.

Benchmarks

Same Qwen3-Coder-30B-A3B-Instruct scaffold, 3-run avg: SWE-bench Verified 39.6% vs 30.4% base (+9.2 pp); SWE-PolyBench Verified 24.5% vs 19.5% (+5.0 pp). Gains persist at 256K with no forced compact and at 16K with shared fallback. Ignoring compact() on the same checkpoint reduces pass rate. Do not invent scores for other bases or vendor products.

Playbook

Read the project page and arXiv HTML. No public weights/standalone repo found—reuse the method (explicit compact tool + execute-vs-ignore ablation), not a product install. Prefer phase-boundary compaction over length-only triggers.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.