autoresearch
v2.2.3
Eval-driven improvement, plus the harness to drive it
Two parts form one loop. An edit → eval → keep/discard engine improves code, prompts, or docs. It scores each change with a shell command, an LLM judge, or both. It shows progress in a live Artifact dashboard. A harness builder measures project health in six categories. It creates feedback loops, evals, sensors, and context management advisories. It then fixes the top-ranked issue through the same loop.
Install
claude plugin install autoresearch@21-breakincode Commands
- /autoresearch:harness-build Menu-driven scaffolder for harness components: feedback loop, eval loop, sensor, or context-mgmt advisory. Writes Tier-1 artifacts into your project's .claude/.
- /autoresearch:harness-check Scan project health across code quality, tests, runtime, architecture, scriptability, and harness completeness. Produces a scored harness report with impact-ranked improvements
- /autoresearch:harness-improvement Execute improvement loop on the top-ranked harness issue: auto-generates eval from probes and spawns the autoresearch:experimenter agent
- /autoresearch:improve Iteratively improve any artifact using an edit-eval-keep/discard loop with live dashboard
Changelog
2.2.3 2026-09-24
- fix state the eval-metrics guard in
/autoresearch:improveonce, with its reason, instead of five times with rising emphasis.
2.2.2 2026-09-19
- feat publish the live improvement dashboard as a private Claude Artifact. The dashboard now uses inline SVG and no external assets.
2.2.1 2026-09-19
- fix rewrite prose across plugin files to clear the simple-english linter. No content or meaning changed, only sentence shape. "harness" stays: standard term for this plugin.
2.2.0 2026-08-10
- feat keep/discard no longer requires git. The experiment loop now snapshots target files into
.autoresearch/snapshot/instead of committing and checking out, so/autoresearch:improveruns against any directory, not just a git repo. - fix stops polluting real repos with per-iteration commits. Nothing read that history. The dashboard sources
reasoninganddiff_summaryfromexperiments.json. - breaking iterations no longer record a commit SHA in
experiments.json(iteration number already identifies them). - test
tests/test_snapshot.sh: 12 assertions, including that a discard reverts to the last kept state rather than the baseline.
2.1.0 2026-08-04
- refactor collapse experiment-loop Rules into steps
- refactor source improve.md libs via ${CLAUDE_PLUGIN_ROOT}
2.0.0 2026-06-28
- feat merge harness plugin into autoresearch (single plugin)
1.2.2 2026-06-27
- feat glass + depth dashboard redesign
1.2.1 2026-06-27
- feat redesign eval dashboard (Linear/Vercel-grade visuals + motion)
1.2.0 2026-06-07
- refactor move harness commands and probes to standalone harness plugin
- refactor convert harness-check and harness-improvement to deprecation shims
1.1.1 2026-06-06
- fix infer metric direction, fix improvement %
1.0.0 2026-05-28
- feat initial release of autoresearch plugin