Eval System
Generate, run, and consciously update evals when AI primitives change — for any repository.
The zoto-eval-system plugin scaffolds two eval backends side by side: a
static backend (pytest, vitest, or jest) for fast, deterministic checks,
and an LLM backend powered by @cursor/sdk for agent-based
evaluations with soft metrics — tokens, duration, verbosity, accuracy, confidence.
LLM cases may be stamped as declarative JSON or code-strategy
TypeScript per target after analyser classification — not as a single global mode.
A diff-aware updater detects when covered targets have changed, classifies each change as critical or non-critical, and presents proposed eval updates for user confirmation before writing anything. Behavioural drift in AI primitives is managed consciously, never silently. User-authored cases stay sovereign — the runtime and compile-time guards both refuse to touch them.
update detects drift after code changes and presents each proposed eval change for your approval.When to use
Reach for the eval system when you want to:
- Stand up an eval suite for a brand-new repo or an existing one, with no bespoke plumbing.
- Detect when eval cases for skills, agents, commands, libraries, and hooks have drifted from the code they describe — and update them with your confirmation.
- Gate CI on drift between code and evals (
pnpm run eval:update --check). - Compare runs across models, dates, or PRs using merged
report.ymldata and Cursor's built-in/canvastool.
The eval system is a coverage and run-quality engine, not a replacement for plugin unit tests. It works alongside any test runner you already have.
Two backends, one suite
Every /z-eval-execute run produces a single timestamped folder under
{evalsDir}/_runs/<ts>/ with three YAML outputs and per-case logs.
| Output | Backend | Role |
|---|---|---|
static.yml |
pytest / vitest / jest | Static-backend totals, per-case pass/fail, framework-specific aggregates. |
llm.yml |
@cursor/sdk |
LLM-backend totals, soft metrics, append-only overlays (drift:, judge:). |
report.yml |
aggregator | Merged rollup with per-case backend: declarative|code. Primary input for /z-eval-compare and CI summaries. Hybrid suites can populate both artifact shapes under one run. |
The LLM backend is always gated on --full +
CURSOR_API_KEY. It never self-runs. CI workflows can opt into LLM evals without
breaking PRs from forks.
User confirmation for every change
No eval is rewritten without your explicit consent. When the updater detects drift or the
configurer proposes a change, you see and approve each modification before it is written.
In CI, the --check mode detects drift without writing — it is your gate, not
an auto-fixer.
Operator approval (configure, update apply) is orthogonal to eval-time
AskQuestion simulation: code-strategy LLM evals replay scripted answers via
evals/llm/_shared/askquestion-bridge.ts without a human in the loop.
See Strategy bridge on the Design page.
Under the hood, commands collect operator answers via askQuestion and relay
them to subagents through a structured needs_user_input / resume contract.
Run /z-eval-init once per repository to drop a fully-commented config template
into .zoto/eval-system/config.yml. Then either edit the YAML by hand or run
/z-eval-configure for the interactive flow.
Strategy bridge
The LLM analyser classifies each covered primitive with
requiresInteraction (and optional interactionStyle). The stamper
routes non-interactive targets to declarative JSON
(plugins/*/evals/**/evals.json) and interactive targets to
code-strategy TypeScript (evals/llm/test_*.test.ts).
Both backends can coexist in one hybrid suite; global llm.strategy in
.zoto/eval-system/config.yml is the fallback default when classification is
unavailable, not the per-target choice.
askquestion-bridge.ts before @cursor/sdk. Full flow on the Design page.The _meta.generated contract
Every generated eval case carries a _meta block. Cases without
_meta, or with _meta.generated === false, are user-authored.
_meta.generated: true. User-authored cases are preserved verbatim — both at runtime and at compile time.- The updater never modifies user-authored cases.
- It flags them for "may need review" and moves on.
- The validator rejects any
update.tssource that does not enforce this guard at compile time (literal-string check for_meta?.generated === true). - Classification metadata (
requiresInteraction,interactionStyle) lives in_meta.primitive_analysison stamped cases — see LLM backend.
See also
_meta.generated contract, run outputs.
→
Configuration
Schema-grounded reference for .zoto/eval-system/config.yml.
→
Source-of-truth: plugins/zoto-eval-system/README.md.