Design
A short tour of the architecture. The lifecycle is six commands wrapped around a
user-confirmed updater. Two backends — static (pytest / vitest / jest) and LLM
(@cursor/sdk) — share the same suite, the same manifest, and the same
_meta.generated contract. Drift in AI primitive behaviour is surfaced
explicitly — every eval change requires your approval before it is written.
Host layout (lean vs ejected)
Greenfield hosts default to lean mode (hostLayout: plugin):
only repo-specific assets live under .zoto/eval-system/ (config, manifest,
cache, nested package.json, scripts/eval-bridge.ts). Engine,
scripts, and templates resolve from the installed plugin at runtime via the bridge and
resolvePluginRoot().
Ejected mode (hostLayout: ejected) is opt-in via
pnpm run eval:stamp-host-layout — it vendors the full runtime into
.zoto/eval-system/ and copies eval primitives to
.cursor/*/eval-sys/. Reverse with pnpm run eval:un-eject
(preserves config, manifest, and eval cases). See the plugin README
Plugin vs host runtime layout section for precedence and migration notes.
Lifecycle architecture
update detects drift after code changes and presents each proposed eval change for your approval before writing.User confirmation for every change
No generated eval is rewritten without explicit user consent. When the updater
detects drift or the configurer proposes a change, every modification is presented
for approval before anything is written to disk. In CI, --check mode
detects drift without writing — it is a gate, not an auto-fixer.
Implementation: the askQuestion / needs_user_input contract
Under the hood, Cursor subagents cannot prompt the user directly. The plumbing works
like this: slash commands own askQuestion calls; they
pre-collect answers before spawning Task subagents and pass them in the Task prompt.
Agents and skills never call askQuestion;
they consume pre-collected answers, or return a structured needs_user_input
block so the command can prompt the user and resume the subagent with the answer.
This is standard Cursor subagent plumbing — the important guarantee is that the user
sees and approves every change. This operator flow is distinct from
eval-runner scripted AskQuestion simulation on code-strategy
cases — see Strategy bridge below. Code-strategy evals
simulate the same contract via
evals/llm/_shared/askquestion-bridge.ts (scripted answers, no human).
Static backend
The default static layout is stamped from templates/static/pytest/ when static.framework is pytest. Vitest and jest variants follow their respective template trees.
conftest.py— shared fixtures.requirements.txt— pytest, pyyaml, and friends.per-primitive-test.py— pattern for generated tests.fixtures/— keepalive directory.
Run it with pnpm run eval (host wiring typically invokes python3 scripts/test.py).
LLM backend (@cursor/sdk)
Why @cursor/sdk: @cursor/sdk ^1.0.12 is the supported, typed
API for building agents and evaluations against Cursor's agent infrastructure. The older
february SDK package is not used anywhere in this plugin.
LLM eval shape is chosen per target by analyser classification — not by a
single repo-wide llm.strategy knob. The
zoto-eval-analyser-subagent emits requiresInteraction (and optional
interactionStyle) into cached analyser payloads; the stamper routes each target
to either declarative JSON or code-strategy TypeScript. Global
llm.strategy in .zoto/eval-system/config.yml is the
fallback default when classification is unavailable; llm.codeFramework
applies to all code-strategy targets.
Strategy bridge
Full pipeline: source primitive → analyser →
{requiresInteraction, interactionStyle} → stamper →
declarative JSON or code-strategy TS →
evals/llm/_shared/askquestion-bridge.ts (interactive code path only) →
@cursor/sdk Agent.run → zoto-llm-reporter /
writer → report.yml row with backend: declarative|code.
AskQuestion turns via askquestion-bridge.ts; each turn is tagged interaction_style: synthetic on SDK 1.0.12.
askquestion-bridge.ts wraps sdk-bridge.ts and replays scripted
answers from case interactions.answers (legacy fallback: follow_ups[]).
Stamped test_*.test.ts files import the bridge through
evals/llm/_shared/; the suite harness in
run-code-strategy-suite.ts wires interactive cases automatically.
| Branch | Stamped artefact | Runner script |
|---|---|---|
requiresInteraction: false | plugins/*/evals/**/evals.json | pnpm run eval:llm:declarative |
requiresInteraction: true | evals/llm/test_*.test.ts | pnpm run eval:llm:code |
Both runners may execute in one pnpm run eval:full pass. Apply-mode update with
reclassification may migrate a target between JSON and TS; eval:cleanup-stale
removes wrong-backend artefacts per target.
The declarative scaffold is stamped at {evalsDir}/_llm/; code-strategy tests live under {evalsDir}/llm/:
| File | Role |
|---|---|
runner.ts | Discovers cases, gates on --full + CURSOR_API_KEY, runs per-case, writes canonical llm.yml + logs/<case>.log. |
case.ts | Typed case loader; understands _meta.generated. |
graders/ | contains, regex, tool-called, llm-judge. |
metrics.ts | Tokens, duration, verbosity, accuracy, confidence. |
writer.ts | Schema-valid YAML + per-case logs. |
update.ts | Surgical diff engine — canonical implementation at plugins/zoto-eval-system/engine/update.ts (stamped mirror under .zoto/eval-system/engine/update.ts). The legacy standalone scripts/eval-update.ts CLI was removed; use pnpm run eval:update instead. |
compare.ts | Emits the /canvas hand-off dataset. |
Model precedence at runtime:
--model <id>on the CLI.ZOTO_EVAL_MODELenv var.config.llm.model.idfrom.zoto/eval-system/config.yml.
Run output layout
Each /z-eval-execute writes a single timestamped folder. Per-case rows in
llm.yml and merged report.yml include
backend: declarative or backend: code so comparers and CI can
group hybrid runs.
{evalsDir}/_runs/<ts>/
static.yml # static-backend totals + per-case rows
llm.yml # LLM-backend totals + soft metrics + per-case rows
report.yml # merged rollup — primary input for compare + CI
.run-meta.json # run id, started_at, model id, git_ref
logs/
<case-id>.log # per-case verbose logs
report.yml is the authoritative aggregate; the comparer reads only report.yml.Result schema (llm.yml example)
schema_version: 1
run_id: 20260503-051900
started_at: 2026-05-03T05:19:00Z
ended_at: 2026-05-03T05:19:42Z
totals:
cases: 14
passed: 12
failed: 2
aggregates:
tokens_total: 48231
duration_ms_total: 42017
verbosity_avg: 0.41
accuracy_avg: 0.86
confidence_avg: 0.72
cases:
- id: zoto-create-spec/1
status: passed
backend: declarative
tokens: 3182
duration_ms: 2710
verbosity: 0.38
accuracy: 0.92
confidence: 0.81
- id: z-eval-configure/2
status: passed
backend: code
interaction_style: synthetic
tokens: 4521
duration_ms: 3890
verbosity: 0.44
accuracy: 0.88
confidence: 0.79
The _meta.generated contract
Every generated eval case carries a _meta block. Cases without _meta, or with _meta.generated === false, are user-authored and are preserved verbatim.
_meta:
generated: true
source_hash: <sha256 of the target's normalised content at generation time>
last_updated: <ISO 8601 timestamp>
generated_by: "zoto-create-evals" | "zoto-update-evals"
primitive_analysis:
requiresInteraction: true
interactionStyle: synthetic
- The updater never modifies user-authored cases.
- It flags them for "may need review" and moves on.
- Classification fields in
_meta.primitive_analysisdrive stamper routing on apply-mode update. - The validator rejects any
update.tssource that does not enforce this guard at compile time (literal-string check for_meta?.generated === true).
Critical-change rubric
| Change | Classification | Reason |
|---|---|---|
| Added target with no eval coverage | critical | New surface area is untested. |
| Removed target with active generated cases | critical | Dangling cases would pass vacuously. |
Skill frontmatter name or description changed | critical | Triggering and retrieval depend on these. |
| Public-surface change on a covered target | critical | The thing being tested behaves differently. |
| Prompt template change | critical | Generated prompts cite the template. |
| Comment-only / whitespace-only change | non-critical | No behaviour change. |
| Internal symbol change (not in public surface) | non-critical | Out of scope for covered behaviour. |
Manifest example
schema_version: 1
created_at: 2026-05-03T05:19:00Z
updated_at: 2026-05-03T05:19:00Z
git_ref: abc1234...def
generated_by: zoto-create-evals
discovery_config:
evalsDir: evals
skillsRoots: [.cursor/skills, skills]
discoveryTargets: [skill, command, agent, hook]
additionalAutomation: []
targets:
- id: skill:zoto-create-spec
kind: skill
path: plugins/zoto-spec-system/skills/zoto-create-spec/SKILL.md
content_hash: 9f8e...a12
public_surface:
frontmatter: { name: zoto-create-spec, description: "Guided..." }
tools: []
eval_files:
- plugins/zoto-spec-system/skills/zoto-create-spec/evals/evals.json
- id: command:z-eval-configure
kind: command
path: plugins/zoto-eval-system/commands/z-eval-configure.md
content_hash: 3a2b...c91
public_surface:
frontmatter: { name: z-eval-configure, description: "Interactive config..." }
tools: []
eval_files:
- evals/llm/test_command_z-eval-configure.test.ts
Judge vs Adviser
The two reviewers complement each other. Judge reads run artefacts after a run; Adviser reads source definitions and eval files before any run.
Adviser (/z-eval-advise) | Judge (/z-eval-judge) | |
|---|---|---|
| When | Before running | After a completed run |
| Reads | Source definitions + eval files | Run artefacts (llm.yml, logs) |
| Answers | "What tests are missing?" | "How did the tests perform?" |
| Output | Gap report + create/update recommendations | Soft-metric annotations in llm.yml |
/z-eval-judge — adversarial review of the latest run. Output lands in llm.yml as an append-only block./z-eval-advise — five-dimension scoreboard plus a deterministic hand-off recommendation.Cross-run comparison
/z-eval-compare <run-1> <run-2> [<run-N>] locates each
run's report.yml, emits a flat dataset (runs × cases × dimensions),
and hands a templates/canvas/compare-prompt.md.tmpl instruction to the host
agent — which invokes Cursor's built-in /canvas tool. The skill never
renders charts itself. The canvas tool decides layout; drilldown opens
{evalsDir}/_runs/<run>/logs/<case>.log.
References
- Source-of-truth:
plugins/zoto-eval-system/README.md. - Schema:
templates/schema/config.schema.json. - Result schema:
templates/schema/result.schema.json.