Eval System

Generate, run, and consciously update evals when AI primitives change — for any repository. The zoto-eval-system plugin scaffolds two eval backends side by side: a static backend (pytest, vitest, or jest) for fast, deterministic checks, and an LLM backend powered by @cursor/sdk for agent-based evaluations with soft metrics — tokens, duration, verbosity, accuracy, confidence. LLM cases may be stamped as declarative JSON or code-strategy TypeScript per target after analyser classification — not as a single global mode.

A diff-aware updater detects when covered targets have changed, classifies each change as critical or non-critical, and presents proposed eval updates for user confirmation before writing anything. Behavioural drift in AI primitives is managed consciously, never silently. User-authored cases stay sovereign — the runtime and compile-time guards both refuse to touch them.

Eval system lifecycle: init → configure → create → execute → judge → compare, with update as a continuous sync loop.
The lifecycle. update detects drift after code changes and presents each proposed eval change for your approval.

When to use

Reach for the eval system when you want to:

  • Stand up an eval suite for a brand-new repo or an existing one, with no bespoke plumbing.
  • Detect when eval cases for skills, agents, commands, libraries, and hooks have drifted from the code they describe — and update them with your confirmation.
  • Gate CI on drift between code and evals (pnpm run eval:update --check).
  • Compare runs across models, dates, or PRs using merged report.yml data and Cursor's built-in /canvas tool.

The eval system is a coverage and run-quality engine, not a replacement for plugin unit tests. It works alongside any test runner you already have.

Two backends, one suite

Every /z-eval-execute run produces a single timestamped folder under {evalsDir}/_runs/<ts>/ with three YAML outputs and per-case logs.

OutputBackendRole
static.yml pytest / vitest / jest Static-backend totals, per-case pass/fail, framework-specific aggregates.
llm.yml @cursor/sdk LLM-backend totals, soft metrics, append-only overlays (drift:, judge:).
report.yml aggregator Merged rollup with per-case backend: declarative|code. Primary input for /z-eval-compare and CI summaries. Hybrid suites can populate both artifact shapes under one run.

The LLM backend is always gated on --full + CURSOR_API_KEY. It never self-runs. CI workflows can opt into LLM evals without breaking PRs from forks.

User confirmation for every change

No eval is rewritten without your explicit consent. When the updater detects drift or the configurer proposes a change, you see and approve each modification before it is written. In CI, the --check mode detects drift without writing — it is your gate, not an auto-fixer.

Operator approval (configure, update apply) is orthogonal to eval-time AskQuestion simulation: code-strategy LLM evals replay scripted answers via evals/llm/_shared/askquestion-bridge.ts without a human in the loop. See Strategy bridge on the Design page.

Under the hood, commands collect operator answers via askQuestion and relay them to subagents through a structured needs_user_input / resume contract. Run /z-eval-init once per repository to drop a fully-commented config template into .zoto/eval-system/config.yml. Then either edit the YAML by hand or run /z-eval-configure for the interactive flow.

Strategy bridge

The LLM analyser classifies each covered primitive with requiresInteraction (and optional interactionStyle). The stamper routes non-interactive targets to declarative JSON (plugins/*/evals/**/evals.json) and interactive targets to code-strategy TypeScript (evals/llm/test_*.test.ts). Both backends can coexist in one hybrid suite; global llm.strategy in .zoto/eval-system/config.yml is the fallback default when classification is unavailable, not the per-target choice.

Eval strategy bridge pipeline: source primitive through analyser classification and stamper routing to declarative JSON or code-strategy TypeScript, with askquestion-bridge on the interactive path, then @cursor/sdk and report.yml with backend annotation.
Analyser-driven per-target routing. Interactive code-strategy cases pass through askquestion-bridge.ts before @cursor/sdk. Full flow on the Design page.

The _meta.generated contract

Every generated eval case carries a _meta block. Cases without _meta, or with _meta.generated === false, are user-authored.

Two columns: generated cases on the left receive update arrows; user-authored cases on the right are preserved verbatim — no arrow touches them.
The updater regenerates only cases marked _meta.generated: true. User-authored cases are preserved verbatim — both at runtime and at compile time.
  • The updater never modifies user-authored cases.
  • It flags them for "may need review" and moves on.
  • The validator rejects any update.ts source that does not enforce this guard at compile time (literal-string check for _meta?.generated === true).
  • Classification metadata (requiresInteraction, interactionStyle) lives in _meta.primitive_analysis on stamped cases — see LLM backend.

See also

Source-of-truth: plugins/zoto-eval-system/README.md.