Design

A short tour of the architecture. The lifecycle is six commands wrapped around a user-confirmed updater. Two backends — static (pytest / vitest / jest) and LLM (@cursor/sdk) — share the same suite, the same manifest, and the same _meta.generated contract. Drift in AI primitive behaviour is surfaced explicitly — every eval change requires your approval before it is written.

Host layout (lean vs ejected)

Greenfield hosts default to lean mode (hostLayout: plugin): only repo-specific assets live under .zoto/eval-system/ (config, manifest, cache, nested package.json, scripts/eval-bridge.ts). Engine, scripts, and templates resolve from the installed plugin at runtime via the bridge and resolvePluginRoot().

Ejected mode (hostLayout: ejected) is opt-in via pnpm run eval:stamp-host-layout — it vendors the full runtime into .zoto/eval-system/ and copies eval primitives to .cursor/*/eval-sys/. Reverse with pnpm run eval:un-eject (preserves config, manifest, and eval cases). See the plugin README Plugin vs host runtime layout section for precedence and migration notes.

Lifecycle architecture

Eval system lifecycle: init, configure, create, execute, judge, compare across the spine, with update sitting below as a continuous sync loop into create and execute.
update detects drift after code changes and presents each proposed eval change for your approval before writing.

User confirmation for every change

No generated eval is rewritten without explicit user consent. When the updater detects drift or the configurer proposes a change, every modification is presented for approval before anything is written to disk. In CI, --check mode detects drift without writing — it is a gate, not an auto-fixer.

Implementation: the askQuestion / needs_user_input contract

Under the hood, Cursor subagents cannot prompt the user directly. The plumbing works like this: slash commands own askQuestion calls; they pre-collect answers before spawning Task subagents and pass them in the Task prompt. Agents and skills never call askQuestion; they consume pre-collected answers, or return a structured needs_user_input block so the command can prompt the user and resume the subagent with the answer. This is standard Cursor subagent plumbing — the important guarantee is that the user sees and approves every change. This operator flow is distinct from eval-runner scripted AskQuestion simulation on code-strategy cases — see Strategy bridge below. Code-strategy evals simulate the same contract via evals/llm/_shared/askquestion-bridge.ts (scripted answers, no human).

Static backend

The default static layout is stamped from templates/static/pytest/ when static.framework is pytest. Vitest and jest variants follow their respective template trees.

  • conftest.py — shared fixtures.
  • requirements.txt — pytest, pyyaml, and friends.
  • per-primitive-test.py — pattern for generated tests.
  • fixtures/ — keepalive directory.

Run it with pnpm run eval (host wiring typically invokes python3 scripts/test.py).

LLM backend (@cursor/sdk)

Why @cursor/sdk: @cursor/sdk ^1.0.12 is the supported, typed API for building agents and evaluations against Cursor's agent infrastructure. The older february SDK package is not used anywhere in this plugin.

LLM eval shape is chosen per target by analyser classification — not by a single repo-wide llm.strategy knob. The zoto-eval-analyser-subagent emits requiresInteraction (and optional interactionStyle) into cached analyser payloads; the stamper routes each target to either declarative JSON or code-strategy TypeScript. Global llm.strategy in .zoto/eval-system/config.yml is the fallback default when classification is unavailable; llm.codeFramework applies to all code-strategy targets.

Strategy bridge

Full pipeline: source primitive → analyser → {requiresInteraction, interactionStyle} → stamper → declarative JSON or code-strategy TS → evals/llm/_shared/askquestion-bridge.ts (interactive code path only) → @cursor/sdk Agent.run → zoto-llm-reporter / writer → report.yml row with backend: declarative|code.

Eval strategy bridge pipeline from source primitive through analyser, classification, stamper, declarative or code-strategy branches, askquestion-bridge on the interactive path, @cursor/sdk, and report.yml with backend annotation.
Per-target routing at stamp time. Code-strategy interactive cases simulate operator AskQuestion turns via askquestion-bridge.ts; each turn is tagged interaction_style: synthetic on SDK 1.0.12.

askquestion-bridge.ts wraps sdk-bridge.ts and replays scripted answers from case interactions.answers (legacy fallback: follow_ups[]). Stamped test_*.test.ts files import the bridge through evals/llm/_shared/; the suite harness in run-code-strategy-suite.ts wires interactive cases automatically.

BranchStamped artefactRunner script
requiresInteraction: falseplugins/*/evals/**/evals.jsonpnpm run eval:llm:declarative
requiresInteraction: trueevals/llm/test_*.test.tspnpm run eval:llm:code

Both runners may execute in one pnpm run eval:full pass. Apply-mode update with reclassification may migrate a target between JSON and TS; eval:cleanup-stale removes wrong-backend artefacts per target.

The declarative scaffold is stamped at {evalsDir}/_llm/; code-strategy tests live under {evalsDir}/llm/:

FileRole
runner.tsDiscovers cases, gates on --full + CURSOR_API_KEY, runs per-case, writes canonical llm.yml + logs/<case>.log.
case.tsTyped case loader; understands _meta.generated.
graders/contains, regex, tool-called, llm-judge.
metrics.tsTokens, duration, verbosity, accuracy, confidence.
writer.tsSchema-valid YAML + per-case logs.
update.tsSurgical diff engine — canonical implementation at plugins/zoto-eval-system/engine/update.ts (stamped mirror under .zoto/eval-system/engine/update.ts). The legacy standalone scripts/eval-update.ts CLI was removed; use pnpm run eval:update instead.
compare.tsEmits the /canvas hand-off dataset.

Model precedence at runtime:

  1. --model <id> on the CLI.
  2. ZOTO_EVAL_MODEL env var.
  3. config.llm.model.id from .zoto/eval-system/config.yml.

Run output layout

Each /z-eval-execute writes a single timestamped folder. Per-case rows in llm.yml and merged report.yml include backend: declarative or backend: code so comparers and CI can group hybrid runs.

{evalsDir}/_runs/<ts>/
  static.yml          # static-backend totals + per-case rows
  llm.yml             # LLM-backend totals + soft metrics + per-case rows
  report.yml          # merged rollup — primary input for compare + CI
  .run-meta.json      # run id, started_at, model id, git_ref
  logs/
    <case-id>.log     # per-case verbose logs
Three-panel run report layout: timestamped run directory on the left, three output files in the centre, /canvas hand-off on the right.
report.yml is the authoritative aggregate; the comparer reads only report.yml.

Result schema (llm.yml example)

schema_version: 1
run_id: 20260503-051900
started_at: 2026-05-03T05:19:00Z
ended_at: 2026-05-03T05:19:42Z
totals:
  cases: 14
  passed: 12
  failed: 2
aggregates:
  tokens_total: 48231
  duration_ms_total: 42017
  verbosity_avg: 0.41
  accuracy_avg: 0.86
  confidence_avg: 0.72
cases:
  - id: zoto-create-spec/1
    status: passed
    backend: declarative
    tokens: 3182
    duration_ms: 2710
    verbosity: 0.38
    accuracy: 0.92
    confidence: 0.81
  - id: z-eval-configure/2
    status: passed
    backend: code
    interaction_style: synthetic
    tokens: 4521
    duration_ms: 3890
    verbosity: 0.44
    accuracy: 0.88
    confidence: 0.79

The _meta.generated contract

Every generated eval case carries a _meta block. Cases without _meta, or with _meta.generated === false, are user-authored and are preserved verbatim.

Two-column diagram. Generated cases on the left receive update arrows; user-authored cases on the right are blocked from any updater write.
Single-owner provenance per case. Drift detection triggers a user-confirmed review, never a silent overwrite.
_meta:
  generated: true
  source_hash: <sha256 of the target's normalised content at generation time>
  last_updated: <ISO 8601 timestamp>
  generated_by: "zoto-create-evals" | "zoto-update-evals"
  primitive_analysis:
    requiresInteraction: true
    interactionStyle: synthetic
  • The updater never modifies user-authored cases.
  • It flags them for "may need review" and moves on.
  • Classification fields in _meta.primitive_analysis drive stamper routing on apply-mode update.
  • The validator rejects any update.ts source that does not enforce this guard at compile time (literal-string check for _meta?.generated === true).

Critical-change rubric

ChangeClassificationReason
Added target with no eval coveragecriticalNew surface area is untested.
Removed target with active generated casescriticalDangling cases would pass vacuously.
Skill frontmatter name or description changedcriticalTriggering and retrieval depend on these.
Public-surface change on a covered targetcriticalThe thing being tested behaves differently.
Prompt template changecriticalGenerated prompts cite the template.
Comment-only / whitespace-only changenon-criticalNo behaviour change.
Internal symbol change (not in public surface)non-criticalOut of scope for covered behaviour.

Manifest example

schema_version: 1
created_at: 2026-05-03T05:19:00Z
updated_at: 2026-05-03T05:19:00Z
git_ref: abc1234...def
generated_by: zoto-create-evals
discovery_config:
  evalsDir: evals
  skillsRoots: [.cursor/skills, skills]
  discoveryTargets: [skill, command, agent, hook]
  additionalAutomation: []
targets:
  - id: skill:zoto-create-spec
    kind: skill
    path: plugins/zoto-spec-system/skills/zoto-create-spec/SKILL.md
    content_hash: 9f8e...a12
    public_surface:
      frontmatter: { name: zoto-create-spec, description: "Guided..." }
      tools: []
    eval_files:
      - plugins/zoto-spec-system/skills/zoto-create-spec/evals/evals.json
  - id: command:z-eval-configure
    kind: command
    path: plugins/zoto-eval-system/commands/z-eval-configure.md
    content_hash: 3a2b...c91
    public_surface:
      frontmatter: { name: z-eval-configure, description: "Interactive config..." }
      tools: []
    eval_files:
      - evals/llm/test_command_z-eval-configure.test.ts

Judge vs Adviser

The two reviewers complement each other. Judge reads run artefacts after a run; Adviser reads source definitions and eval files before any run.

Adviser (/z-eval-advise)Judge (/z-eval-judge)
WhenBefore runningAfter a completed run
ReadsSource definitions + eval filesRun artefacts (llm.yml, logs)
Answers"What tests are missing?""How did the tests perform?"
OutputGap report + create/update recommendationsSoft-metric annotations in llm.yml
Cursor IDE mockup of /z-eval-judge. The chat panel summarises a weak grader, a verbosity spike, and an accuracy regression; the editor pane shows the judge: block being appended to llm.yml.
/z-eval-judge — adversarial review of the latest run. Output lands in llm.yml as an append-only block.
Cursor IDE mockup of /z-eval-advise. The right pane shows a five-dimension severity scoreboard (trigger-phrase, schema, regression, citation, checklist) and a hand-off banner recommending /z-eval-update.
/z-eval-advise — five-dimension scoreboard plus a deterministic hand-off recommendation.

Cross-run comparison

/z-eval-compare <run-1> <run-2> [<run-N>] locates each run's report.yml, emits a flat dataset (runs × cases × dimensions), and hands a templates/canvas/compare-prompt.md.tmpl instruction to the host agent — which invokes Cursor's built-in /canvas tool. The skill never renders charts itself. The canvas tool decides layout; drilldown opens {evalsDir}/_runs/<run>/logs/<case>.log.

Cursor IDE mockup of /z-eval-compare run-A run-B with a /canvas chart of runs × cases × tokens for two run series.
Click any bar to open the per-case log in the editor.

References