Skip to content

Repository evaluation tutorial ​

This optional tutorial is for contributors working in a clone of the Intention Kernel repository. If you installed the package in an application, start with Evaluating from your project; no repository scripts are required.

This tutorial runs a complete declarative suite, saves its report, and demonstrates a business failure that simpler structural checks would miss. It uses the public intention-kernel/testing API and requires no model credentials.

Run the example ​

From the repository root, with Node.js 22 and npm 10:

sh
npm ci
npm run demo:evaluation

If dependencies are already installed, run only the second command. The script builds the package, executes examples/evaluation.ts, and loads examples/evaluation.suite.json.

Expected result: passed: true, four passed cases, and a reportFile path under examples/.runs/evaluation-<runId>/report.json. Each invocation creates a new report directory by default.

The adapter is a small, scripted catalogue. It recognizes the exact messages in the JSON fixture and returns predictable observations. This verifies the example's evaluation logic; evaluating model understanding requires connecting a real agent.

Read the scenario definitions ​

The suite expresses two business scenarios:

ScenarioInputs and casesAcceptance criteria
choose-productTwo ways to request products, followed by selecting the first option. Two cases, each with two turns.The search completes, offers a choice with two options, persists a selection, closes the choice, and stores the exact product requested by the user.
availabilitySearch a catalogue with results or an empty catalogue. Two cases, each with one turn.The search completes and either offers a choice with results, or reports no results with no choice.

That is four cases and six turns. The final correct-product evaluator compares the requested option from the search turn with the stored selection after the second turn.

Complete evaluation.suite.json
json
{
  "schemaVersion": 1,
  "id": "catalogue",
  "name": "Catalogue conversations",
  "scenarios": [
    {
      "id": "choose-product",
      "name": "Choose a product from the displayed options",
      "description": "The chosen product must be the option requested by the user, not just any valid product.",
      "tags": ["catalogue", "selection"],
      "writePolicy": "read_only",
      "steps": [
        {
          "id": "search",
          "input": { "text": "Show available products" },
          "variants": [
            { "id": "direct", "input": { "text": "Show available products" } },
            { "id": "polite", "input": { "text": "Please show available products" } }
          ],
          "assertions": [
            { "id": "response-completed", "kind": "check", "path": ["response", "status"], "operator": "equals", "value": "completed" },
            { "id": "choice-offered", "kind": "check", "path": ["response", "interaction", "kind"], "operator": "equals", "value": "choice" },
            { "id": "two-options", "kind": "check", "path": ["response", "interaction", "options"], "operator": "length", "value": 2 }
          ]
        },
        {
          "id": "select",
          "input": { "text": "I choose the first product", "selection": { "optionIndex": 0 } },
          "assertions": [
            { "id": "selection-persisted", "kind": "check", "path": ["checkpoint", "facts"], "operator": "contains", "value": { "type": "product.selected" } },
            { "id": "choice-closed", "kind": "check", "path": ["response", "interaction"], "operator": "absent" }
          ]
        }
      ],
      "assertions": [
        {
          "id": "correct-product",
          "label": "The stored selection matches the requested option from the search turn",
          "kind": "custom",
          "evaluator": "catalogue.selected-option",
          "parameters": { "searchStepId": "search", "selectionStepId": "select" }
        }
      ]
    },
    {
      "id": "availability",
      "name": "Report the available catalogue outcome",
      "tags": ["catalogue", "alternatives"],
      "writePolicy": "read_only",
      "steps": [
        {
          "id": "search",
          "input": { "text": "Show available products" },
          "variants": [
            { "id": "available", "input": { "text": "Show available products" } },
            { "id": "empty", "input": { "text": "Show unavailable products" } }
          ],
          "assertions": [
            { "id": "response-completed", "kind": "check", "path": ["response", "status"], "operator": "equals", "value": "completed" },
            {
              "id": "accepted-outcome",
              "kind": "any",
              "assertions": [
                {
                  "id": "products-offered",
                  "kind": "all",
                  "assertions": [
                    { "id": "has-results", "kind": "check", "path": ["catalogue", "count"], "operator": "gte", "value": 1 },
                    { "id": "has-choice", "kind": "check", "path": ["response", "interaction", "kind"], "operator": "equals", "value": "choice" }
                  ]
                },
                {
                  "id": "empty-catalogue-explained",
                  "kind": "all",
                  "assertions": [
                    { "id": "no-results", "kind": "check", "path": ["catalogue", "count"], "operator": "equals", "value": 0 },
                    { "id": "no-choice", "kind": "check", "path": ["response", "interaction"], "operator": "absent" },
                    { "id": "reason-published", "kind": "check", "path": ["response", "reason"], "operator": "equals", "value": "no_results" }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "assertions": [
    { "id": "no-pending-cases", "kind": "check", "path": ["summary", "pending"], "operator": "equals", "value": 0 }
  ]
}

The file above is included directly from the executable fixture. Edit that file to change the tutorial's definitions.

catalogue.count, response.reason, and the fact type product.selected belong to this example's adapter and domain. They are not fields or business rules supplied automatically by the framework. Observation paths depend on the adapter you use.

Introduce a failure ​

sh
npm run demo:evaluation -- --fail

The script now deliberately stores the second product when the user selects the first. Expected result: two failed cases, two passed cases, passed: false, and exit code 1. The nonzero exit is intentional in this exercise.

The response still completes, a selection fact exists, and the choice closes. Those checks pass. The final business criterion catches the incorrect product. In the failed case's assertions, the relevant report excerpt is:

json
{
  "id": "correct-product",
  "label": "The stored selection matches the requested option from the search turn",
  "passed": false,
  "expected": {
    "evaluator": "catalogue.selected-option",
    "parameters": { "searchStepId": "search", "selectionStepId": "select" }
  },
  "evidence": {
    "searchStepId": "search",
    "selectionStepId": "select",
    "expectedProductId": "desk-lamp",
    "actualProductId": "floor-lamp"
  }
}

This is why a scenario can need several acceptance criteria: the presence of a fact does not prove that its value matches the user's request.

Select cases and repeat them ​

sh
npm run demo:evaluation -- --scenario=choose-product
npm run demo:evaluation -- --scenario=choose-product --repetitions=3

The first command runs the two selection cases. The second runs six cases, each with a separate session. Repetitions of this scripted fixture remain deterministic; repetitions against a real model measure repeated attempts under the configured model and environment.

The tutorial script accepts these arguments:

ArgumentEffect
--suite=path/to/suite.jsonLoad another suite; paths are relative to the current working directory. It must use this adapter's inputs and registered evaluators.
--scenario=choose-productRun one scenario ID, including its variants.
--repetitions=3Repeat each expanded case three times.
--output=examples/.runs/my-report.jsonWrite the complete report to this path, replacing an existing file at that path.
--failEnable the intentional selection bug in this tutorial's adapter.

These arguments belong to examples/evaluation.ts. The library exposes functions; it does not install a universal evaluation CLI.

How the script executes the suite ​

The execution and report-writing section is shown below. Its argument() helper, adapter, evaluators, and simulateFailure flag are defined earlier in examples/evaluation.ts.

ts
const suiteFile = argument("suite") ?? fileURLToPath(new URL("./evaluation.suite.json", import.meta.url));
const suite = parseEvaluationSuite(JSON.parse(await readFile(suiteFile, "utf8")) as unknown);
const scenarioId = argument("scenario");
const controller = new AbortController();
const cancel = (): void => { controller.abort(); };
process.once("SIGINT", cancel);
process.once("SIGTERM", cancel);

const report = await runEvaluation({
  suite, adapter, evaluators,
  target: { id: "tutorial.catalogue", fingerprint: simulateFailure ? "wrong-selection" : "correct-selection" },
  ...(scenarioId === undefined ? {} : { scenarioIds: [scenarioId] }),
  concurrency: 2,
  repetitions: Number(argument("repetitions") ?? 1),
  signal: controller.signal,
}).finally(() => {
  process.removeListener("SIGINT", cancel);
  process.removeListener("SIGTERM", cancel);
});

const output = path.resolve(argument("output") ?? path.join(
  "examples", ".runs", `evaluation-${report.runId}`, "report.json",
));
await mkdir(path.dirname(output), { recursive: true });
await writeFile(output, JSON.stringify(report, null, 2) + "\n");

// This tutorial requires every selected case and every suite assertion to pass.
// A case already aggregates its complete assertion expression, including any/all.
const passed = report.status === "completed"
  && report.cases.every((entry) => entry.status === "passed")
  && (report.assertions ?? []).every((assertion) => assertion.passed);
process.stdout.write(JSON.stringify({ passed, summary: report.summary, reportFile: output }, null, 2) + "\n");
if (!passed) process.exitCode = 1;

The example's exit policy requires a completed run, every selected case to pass, and every suite assertion to pass. The runner itself returns a report; your host chooses how to surface that result in a command or CI job.

Apply this to your agent ​

  1. Describe the use case as one or more scenarios, with all required turn and final criteria.
  2. Connect your compiled agent with createAgentEvaluationAdapter(), or implement an adapter for your application's conversation API.
  3. Add custom evaluators for relationships that cannot be checked against one observation.
  4. Preserve the returned evidence and use the report-driven iteration workflow.

The existing Gemini examples show model-backed conversations using the same runner, with per-case kernels, stores, and simulated write ports.