Quickstart
From a fresh Cursor install to a compared eval run in under fifteen minutes. Each step maps to one slash command; the lifecycle is the same on Linux, macOS, and Windows.
Install
Add the plugin to Cursor (marketplace install or local dev install from this monorepo):
cursor plugin install zoto-eval-system
No host-repo files change at this point. The plugin fails loudly if any command other
than /z-eval-init runs before configuration exists.
Init — drop the config template
One-time scaffold. Writes a fully-commented template to
.zoto/eval-system/config.yml with every key showing the internal default the
plugin would otherwise apply.
/z-eval-init
Init also stamps the lean eval-home runtime under
.zoto/eval-system/ (nested package.json,
scripts/eval-bridge.ts, cache dirs). Run orchestrated evals with
pnpm -C .zoto/eval-system run eval / eval:full once
dependencies are installed. Edit config by hand if you prefer — every line is
commented, so you can uncomment selectively — or run /z-eval-configure
for the interactive flow.
Configure — interactive guided setup
The command walks you through each config field interactively, validates your choices,
emits a cleanup_plan when framework or strategy changes, and writes
.zoto/eval-system/config.yml in one atomic write.
/z-eval-configure
Key prompts: static.framework, llm.strategy (sets the default fallback only), llm.codeFramework, evalsDir, discoveryTargets[], llm.runtime, llm.model.id, judgeModel. After /z-eval-create, the analyser may stamp a hybrid mix of declarative JSON and code-strategy TS per target.
What gets stamped where?
Non-interactive targets → plugins/*/evals/**/evals.json (declarative).
Interactive targets → evals/llm/test_*.test.ts (code-strategy + bridge).
See the Strategy bridge section on the Design page
for the full analyser → stamper → backend pipeline.
Create — stamp the suite
Discovers covered targets (skills, commands, agents, hooks), runs the LLM analyser per
approved central primitive, stamps both static and LLM backends, and writes
manifest.yml + manifest.history.yml.
/z-eval-create
What lands on disk:
evals/— pytest backend (or vitest/jest if configured).evals/_llm/— declarative@cursor/sdkrunner, graders, writer, update engine.evals/llm/test_*.test.ts— code-strategy LLM evals for interactive targets (analyser setsrequiresInteraction).plugins/*/evals/**/evals.json— declarative LLM cases for non-interactive targets..zoto/eval-system/manifest.yml— covered targets + discovery snapshot..zoto/eval-system/manifest.history.yml— append-only history..env.example— placeholder forCURSOR_API_KEY(only created if absent).
Update — review and approve eval changes after code drift
Run after any edit to a covered target (skill body, agent body, command, library export,
hook script). Or run in CI as pnpm run eval:update --check for a drift gate.
| Invocation | Mode | Interactive? | Writes? |
|---|---|---|---|
/z-eval-update | rediscovery dry-run | no | no |
/z-eval-update --apply | rediscovery apply | yes (askQuestion per change) | yes |
/z-eval-update <file-or-glob> | targeted dry-run | no | no |
/z-eval-update <file-or-glob> --apply | targeted apply | yes | yes |
pnpm run eval:update --check | CI check | no | no |
Rediscovery uses the snapshot in manifest.discovery_config, not the current
config.yml — config edits never masquerade as code drift. Apply-mode may
migrate a target between declarative JSON and code-strategy TS when analyser
classification flips; the configurer's cleanup_plan removes stale artefacts.
Execute — capture a baseline
The orchestrator invokes the host package.json scripts and writes the run
directory under {evalsDir}/_runs/<ts>/.
/z-eval-execute # static only — fast, no API key required
/z-eval-execute --full # static + LLM — gated on CURSOR_API_KEY
What lands on disk per run:
evals/_runs/20260506T080011Z/
static.yml # pytest totals + per-case rows
llm.yml # LLM totals + soft metrics + per-case rows
report.yml # aggregate of static + llm (per-case backend: declarative|code)
.run-meta.json # run id, started_at, model id, git_ref
logs/<case-id>.log # per-case verbose logs
report.yml is the authoritative aggregate.Judge — adversarial review of the latest run
The judge agent reads static.yml, llm.yml, the merged
report.yml, and per-case logs, then appends judge: annotations
into llm.yml.
/z-eval-judge
What it surfaces:
- Weak graders —
containswhereregexwould be stricter. - Verbosity spikes — token totals far above the per-case baseline.
- Accuracy regressions — soft-metric drops vs the previous run.
The command (not the skill) offers, via askQuestion, to run
/z-eval-update on the affected skills. The judge subagent never calls
askQuestion directly.
Advise — pre-hoc coverage gap analysis
Where Judge tells you how a run performed, Advise tells you what tests are missing before you run anything. Reads source definitions and eval files only — never run artefacts — and scores five dimensions:
/z-eval-advise # full scan
/z-eval-advise <plugin-name> # scope to one plugin
/z-eval-advise <skill-name> # scope to one skill
Two interactive breakpoints — drill-down on per-dimension severity, then accept / walk /
skip the recommendations. Accepted recommendations hand off to
/z-eval-create (no eval file exists) or /z-eval-update --target <glob> --apply.
The adviser never modifies files itself.
Compare — cross-run on /canvas
Pass two or more run ids. The comparer locates each run's report.yml, emits a
flat dataset (runs × cases × dimensions, including a backend
dimension for hybrid suites), and hands a rendering instruction
to the host agent — which invokes Cursor's built-in /canvas tool. The skill
never renders charts itself.
/z-eval-compare run-A run-B
{evalsDir}/_runs/<ts>/logs/<case-id>.log.CI integration
A drop-in GitHub Actions gate. Critical drift exits the workflow with code 2; LLM evals
stay opt-in via the CURSOR_API_KEY secret.
- run: pnpm install
- run: pnpm run eval:update --check # exits 2 on critical drift
- run: pnpm run eval # static
- if: env.CURSOR_API_KEY != ''
run: pnpm run eval:full # LLM — gated on secret
env:
CURSOR_API_KEY: ${{ secrets.CURSOR_API_KEY }}