Quickstart

From a fresh Cursor install to a compared eval run in under fifteen minutes. Each step maps to one slash command; the lifecycle is the same on Linux, macOS, and Windows.

Install

Add the plugin to Cursor (marketplace install or local dev install from this monorepo):

cursor plugin install zoto-eval-system

No host-repo files change at this point. The plugin fails loudly if any command other than /z-eval-init runs before configuration exists.

Init — drop the config template

One-time scaffold. Writes a fully-commented template to .zoto/eval-system/config.yml with every key showing the internal default the plugin would otherwise apply.

/z-eval-init

Init also stamps the lean eval-home runtime under .zoto/eval-system/ (nested package.json, scripts/eval-bridge.ts, cache dirs). Run orchestrated evals with pnpm -C .zoto/eval-system run eval / eval:full once dependencies are installed. Edit config by hand if you prefer — every line is commented, so you can uncomment selectively — or run /z-eval-configure for the interactive flow.

Configure — interactive guided setup

The command walks you through each config field interactively, validates your choices, emits a cleanup_plan when framework or strategy changes, and writes .zoto/eval-system/config.yml in one atomic write.

/z-eval-configure

Key prompts: static.framework, llm.strategy (sets the default fallback only), llm.codeFramework, evalsDir, discoveryTargets[], llm.runtime, llm.model.id, judgeModel. After /z-eval-create, the analyser may stamp a hybrid mix of declarative JSON and code-strategy TS per target.

What gets stamped where?

Non-interactive targets → plugins/*/evals/**/evals.json (declarative). Interactive targets → evals/llm/test_*.test.ts (code-strategy + bridge). See the Strategy bridge section on the Design page for the full analyser → stamper → backend pipeline.

Create — stamp the suite

Discovers covered targets (skills, commands, agents, hooks), runs the LLM analyser per approved central primitive, stamps both static and LLM backends, and writes manifest.yml + manifest.history.yml.

/z-eval-create

What lands on disk:

  • evals/ — pytest backend (or vitest/jest if configured).
  • evals/_llm/ — declarative @cursor/sdk runner, graders, writer, update engine.
  • evals/llm/test_*.test.ts — code-strategy LLM evals for interactive targets (analyser sets requiresInteraction).
  • plugins/*/evals/**/evals.json — declarative LLM cases for non-interactive targets.
  • .zoto/eval-system/manifest.yml — covered targets + discovery snapshot.
  • .zoto/eval-system/manifest.history.yml — append-only history.
  • .env.example — placeholder for CURSOR_API_KEY (only created if absent).

Update — review and approve eval changes after code drift

Run after any edit to a covered target (skill body, agent body, command, library export, hook script). Or run in CI as pnpm run eval:update --check for a drift gate.

InvocationModeInteractive?Writes?
/z-eval-updaterediscovery dry-runnono
/z-eval-update --applyrediscovery applyyes (askQuestion per change)yes
/z-eval-update <file-or-glob>targeted dry-runnono
/z-eval-update <file-or-glob> --applytargeted applyyesyes
pnpm run eval:update --checkCI checknono

Rediscovery uses the snapshot in manifest.discovery_config, not the current config.yml — config edits never masquerade as code drift. Apply-mode may migrate a target between declarative JSON and code-strategy TS when analyser classification flips; the configurer's cleanup_plan removes stale artefacts.

Execute — capture a baseline

The orchestrator invokes the host package.json scripts and writes the run directory under {evalsDir}/_runs/<ts>/.

/z-eval-execute            # static only — fast, no API key required
/z-eval-execute --full     # static + LLM — gated on CURSOR_API_KEY

What lands on disk per run:

evals/_runs/20260506T080011Z/
  static.yml          # pytest totals + per-case rows
  llm.yml             # LLM totals + soft metrics + per-case rows
  report.yml          # aggregate of static + llm (per-case backend: declarative|code)
  .run-meta.json      # run id, started_at, model id, git_ref
  logs/<case-id>.log  # per-case verbose logs
Three-panel run-report layout: timestamped run directory on the left, three output files in the middle, /canvas hand-off on the right.
Each run produces three YAML outputs. report.yml is the authoritative aggregate.

Judge — adversarial review of the latest run

The judge agent reads static.yml, llm.yml, the merged report.yml, and per-case logs, then appends judge: annotations into llm.yml.

/z-eval-judge

What it surfaces:

  • Weak graders — contains where regex would be stricter.
  • Verbosity spikes — token totals far above the per-case baseline.
  • Accuracy regressions — soft-metric drops vs the previous run.

The command (not the skill) offers, via askQuestion, to run /z-eval-update on the affected skills. The judge subagent never calls askQuestion directly.

Advise — pre-hoc coverage gap analysis

Where Judge tells you how a run performed, Advise tells you what tests are missing before you run anything. Reads source definitions and eval files only — never run artefacts — and scores five dimensions:

/z-eval-advise                      # full scan
/z-eval-advise <plugin-name>        # scope to one plugin
/z-eval-advise <skill-name>         # scope to one skill

Two interactive breakpoints — drill-down on per-dimension severity, then accept / walk / skip the recommendations. Accepted recommendations hand off to /z-eval-create (no eval file exists) or /z-eval-update --target <glob> --apply. The adviser never modifies files itself.

Compare — cross-run on /canvas

Pass two or more run ids. The comparer locates each run's report.yml, emits a flat dataset (runs × cases × dimensions, including a backend dimension for hybrid suites), and hands a rendering instruction to the host agent — which invokes Cursor's built-in /canvas tool. The skill never renders charts itself.

/z-eval-compare run-A run-B
Cursor IDE mockup: /z-eval-compare run-A run-B rendering a /canvas chart with two run series.
Drilldown opens the per-case log under {evalsDir}/_runs/<ts>/logs/<case-id>.log.

CI integration

A drop-in GitHub Actions gate. Critical drift exits the workflow with code 2; LLM evals stay opt-in via the CURSOR_API_KEY secret.

- run: pnpm install
- run: pnpm run eval:update --check   # exits 2 on critical drift
- run: pnpm run eval                  # static
- if: env.CURSOR_API_KEY != ''
  run: pnpm run eval:full             # LLM — gated on secret
  env:
    CURSOR_API_KEY: ${{ secrets.CURSOR_API_KEY }}

Next steps