Skip to content

Evaluation framework ​

Write a conversation as data: the inputs to send, the behavior to accept after each turn, and the conditions that must hold at the end. The evaluation runner executes that definition and produces structured results for every criterion, with the observations needed to investigate failures.

A scenario can require many acceptance criteria. For example, choosing a product can require valid options, the correct selection, a persisted fact, and a closed interaction. Use multiple assertions for simultaneous requirements, any for acceptable alternatives, and custom evaluators for relationships across turns.

Import the framework from intention-kernel/testing. It is independent of the model provider and can run through a compiled agent, an HTTP adapter, a CLI, or a host application.

Start here ​

GoalGuide
Evaluate an agent from your installed packageEvaluating from your project
Translate a use case into scenarios, steps, inputs, and variantsWriting scenarios
Express several criteria, alternatives, and business rulesAssertions and evaluators
Connect your agent and control executionRunning evaluations
Read the evidence and let a coding agent iterateReports and agent iteration

Use the installed library ​

Call the public API from your application's code:

ts
import type { CompiledAgent } from "intention-kernel";
import {
  createAgentEvaluationAdapter,
  parseEvaluationSuite,
  runEvaluation,
  type EvaluationEvaluators,
} from "intention-kernel/testing";

export function evaluateScenarios(
  agent: CompiledAgent,
  definition: unknown,
  evaluators: EvaluationEvaluators = {},
) {
  return runEvaluation({
    suite: parseEvaluationSuite(definition),
    target: { id: agent.id, fingerprint: agent.fingerprint },
    adapter: createAgentEvaluationAdapter({ agent }),
    evaluators,
  });
}

The wrapper above belongs to your application. Pass its existing compiled agent and a declarative suite; it returns the report to the caller. No repository checkout, build, npm script, or specific test framework is required. The first evaluation supplies a complete suite and shows how to use its results.

The definition model ​

TermMeaning
Use caseA business objective you describe through one or more scenarios. There is no separate useCase object in the schema.
SuiteA versioned collection of scenarios, with optional checks across cases.
ScenarioAn ordered conversation and its acceptance criteria.
StepOne input sent to the adapter and the checks on its resulting observation.
VariantA complete alternative input for one step.
CaseOne scenario, one combination of step variants, and one repetition. It receives its own conversation identity.
ObservationThe JSON evidence returned by the adapter for a turn.
AssertionA declared criterion evaluated against that evidence.
Custom evaluatorHost code registered under a name to check a domain relationship or semantic requirement.
ReportThe definition snapshot, case inventory, submitted inputs, observations, assertion results, and execution status.

How validation works ​

For each selected case, the runner opens a session and sends the steps in order. It evaluates each step's assertions after receiving its observation. Ordinary assertion failures remain in the report and do not prevent later steps from running; transport or evaluator errors can interrupt the conversation.

After the defined steps have been recorded and the final observation is available, scenario assertions evaluate the final state. Custom scenario evaluators also receive the complete recorded turn history. Suite assertions run after scheduling ends and can inspect the whole selected case inventory.

At each scope, all assertions in the array are required. Inside an assertion, all, any, and not express the acceptance logic. The report preserves that expression and its child results, so a failed alternative inside a passing any is distinguishable from a failed required criterion.

Completion is not acceptance

report.status === "completed" means execution finished. Cases may still have failed, errored, been skipped, or been blocked. Suite assertions have separate results in report.assertions. Read the report contract before turning it into an exit code or CI result.

What you supply ​

The framework supplies definition validation, case expansion, scheduling, assertion composition, and portable reports. Your application supplies:

  • the agent or conversation transport;
  • fixture setup and the state isolation needed by the business case;
  • custom evaluators and their diagnostic evidence;
  • model/provider configuration when evaluating a real agent;
  • the application entry point that calls the API and any report storage or presentation it needs.

The JSON definition contains data. It does not execute JavaScript or automatically interpret business descriptions as assertions. A coding agent can edit the definition, invoke your application's evaluation entry point, inspect the returned report, change the implementation, and repeat the same scenarios.

For exact exported signatures, see the testing API reference.