Anthropic's Agent Skills (packaged into SKILL.md directories) encapsulate procedural know-how for LLM agents, with open aggregations expanding beyond 230,000 skills. While prior literature relied on expensive LLM-mediated loops within the agent decision process, SkillSeek introduces a two-stage IR pipeline (BGE-base bi-encoder + small cross-encoder) exposed via Model Context Protocol (MCP). Evaluated across 89 SkillsBench tasks on OpenHands, SkillSeek matches or exceeds LLM-mediated retrieval loops while slashing per-trial spend from $51.30 to $27.54 (a 46.3% reduction).

Key Takeaways

  • ✓Slashes per-trial retrieval spend from $51.30 to $27.54, reducing token costs by 46.3%
  • ✓Matches costly LLM-mediated loops across 89 SkillsBench tasks via BGE-base bi-encoders and compact cross-encoders
  • ✓Native Model Context Protocol (MCP) server support enables seamless drop-in integration with OpenHands, Claude Code, and agent harnesses
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

核心背景与行业痛点 With Anthropic standardizing procedural agent know-how into SKILL.md directories, open-source aggregations have rapidly ballooned past 230,000 skills. However, selecting the right skill has replaced authoring as the central bottleneck. The prevailing practice outsources selection to the agent's internal LLM decision loop, continuously rewriting queries and scoring candidates, which inflicts massive token overhead and severe latency. ### 架构亮点与底层机制 SkillSeek re-architects agent skill discovery by leveraging standard Information Retrieval (IR) pipelines exposed over the Model Context Protocol (MCP). It features a two-stage retrieval pipeline: a BGE-base bi-encoder paired with BM25 for rapid candidate pruning to Top-50, followed by a compact cross-encoder for fine-grained ranking. By serving over MCP, any compatible agent runtime (such as Claude Code, OpenHands, Cline, or Antigravity) retrieves skills via standard deterministic tool calls without paying continuous model introspection tokens. ### 权威 Benchmark 与实测跑分对比 Evaluated on the 89-task SkillsBench benchmark under the OpenHands harness across a 4x11 grid of pools, backbones, and retrieval methods: 1. Parity with LLM Loops: Plain BM25 matches or exceeds expensive LLM-mediated loops on 3 of 4 settings, and the compact cross-encoder bridges the remaining gap, achieving full performance parity. 2. Cost Slashed by 46%: Average per-trial spend falls from $51.30 under the LLM loop down to $27.54 with SkillSeek—within $0.50 of the no-skill baseline—achieving near-zero marginal inference overhead for skill retrieval. ### 开发者实战落地与开箱指南 SkillSeek is fully open-sourced on GitHub with ready-to-run Docker images and lightweight Python packages. Developers configuring MCP clients can register SkillSeek under their agent settings, allowing agents to dynamically query the global skill registry via deterministic MCP tool calls without manual skill subset filtering.