autoresearch
v2.2.3
以評測驅動改進,並提供驅動它的 harness
外掛由兩個部分組成一個循環。核心是「編輯 → 評測 → 保留或捨棄」引擎。引擎可以改進程式碼、提示詞或文件。引擎用 shell 指令、LLM 評審或兩者,為每次修改評分。進度顯示在即時的 Artifact 儀表板上。Harness 建構器從六個類別衡量專案健康度。建構器會建立回饋循環、評測、感測器,以及管理脈絡的建議。然後,建構器用同一個循環修正排名最高的問題。
安裝
claude plugin install autoresearch@21-breakincode 指令
- /autoresearch:harness-build Menu-driven scaffolder for harness components: feedback loop, eval loop, sensor, or context-mgmt advisory. Writes Tier-1 artifacts into your project's .claude/.
- /autoresearch:harness-check Scan project health across code quality, tests, runtime, architecture, scriptability, and harness completeness. Produces a scored harness report with impact-ranked improvements
- /autoresearch:harness-improvement Execute improvement loop on the top-ranked harness issue: auto-generates eval from probes and spawns the autoresearch:experimenter agent
- /autoresearch:improve Iteratively improve any artifact using an edit-eval-keep/discard loop with live dashboard
更新紀錄
2.2.3 2026-09-24
- fix state the eval-metrics guard in
/autoresearch:improveonce, with its reason, instead of five times with rising emphasis.
2.2.2 2026-09-19
- feat publish the live improvement dashboard as a private Claude Artifact. The dashboard now uses inline SVG and no external assets.
2.2.1 2026-09-19
- fix rewrite prose across plugin files to clear the simple-english linter. No content or meaning changed, only sentence shape. "harness" stays: standard term for this plugin.
2.2.0 2026-08-10
- feat keep/discard no longer requires git. The experiment loop now snapshots target files into
.autoresearch/snapshot/instead of committing and checking out, so/autoresearch:improveruns against any directory, not just a git repo. - fix stops polluting real repos with per-iteration commits. Nothing read that history. The dashboard sources
reasoninganddiff_summaryfromexperiments.json. - breaking iterations no longer record a commit SHA in
experiments.json(iteration number already identifies them). - test
tests/test_snapshot.sh: 12 assertions, including that a discard reverts to the last kept state rather than the baseline.
2.1.0 2026-08-04
- refactor collapse experiment-loop Rules into steps
- refactor source improve.md libs via ${CLAUDE_PLUGIN_ROOT}
2.0.0 2026-06-28
- feat merge harness plugin into autoresearch (single plugin)
1.2.2 2026-06-27
- feat glass + depth dashboard redesign
1.2.1 2026-06-27
- feat redesign eval dashboard (Linear/Vercel-grade visuals + motion)
1.2.0 2026-06-07
- refactor move harness commands and probes to standalone harness plugin
- refactor convert harness-check and harness-improvement to deprecation shims
1.1.1 2026-06-06
- fix infer metric direction, fix improvement %
1.0.0 2026-05-28
- feat initial release of autoresearch plugin