AI checks
Ask a model you choose to judge text, screenshots or recorded frames against a requirement you wrote.
An AI check asks a judge, a model you choose, whether evidence meets a requirement you wrote. Use one where the requirement needs reading, such as whether an error message says how to go on. Check facts with expect.
test('explains the saved task', async ({ page }) => { await page.goto('/') await page.getByTestId('task-title').fill('Release checklist') await page.getByTestId('save-task').click() await expect(page.getByTestId('saved-task')).toHaveText('Release checklist') await test.evaluate({ requirement: 'The message says the task was saved and names it.', evidence: { capture: 'screenshot' }, })})Retest's own process takes the screenshot, calls the judge, checks the answer and records the verdict. The test's process only names what to judge. It never holds the judge's key and never loads its code.
Judges Link to Judges
The config declares its judges under evaluation. A run of ordinary tests loads no judge and reads no key.
evaluation: { judges: { visual: { adapter: '@rehearsal-labs/retest/evaluation/ai-sdk', credentials: { apiKey: env('ANTHROPIC_API_KEY') }, options: { provider: 'anthropic', model: '<model-id>' }, accepts: ['text', 'images'], }, }, defaultJudge: 'visual', timeoutMs: 30_000, limits: { callsPerTest: 5, callsPerRun: 100 },},| Key | What it holds |
|---|---|
adapter | What makes the judge: a package, a module of your own, or a function |
credentials | The judge's keys, from env(name) or a function. A key written as its value is refused. |
options | The judge's settings, as JSON values |
accepts | What the judge takes: text, images or frames. It gets nothing else. |
defaultJudge | The judge a check uses when it names none |
timeoutMs | How long a check may take. The default is 30 seconds. |
limits | Bounds on calls, images, frames and output for the whole run |
Providers Link to Providers
The adapter at @rehearsal-labs/retest/evaluation/ai-sdk judges through the Vercel AI SDK. Retest does not install the SDK, so add ai@^7.0.127 and the package for your provider:
| Provider | Package | Options |
|---|---|---|
| Anthropic | @ai-sdk/anthropic@^4.0.71 | provider: 'anthropic', model |
| OpenAI | @ai-sdk/openai@^4.0.83 | provider: 'openai', model |
| Azure OpenAI | @ai-sdk/azure@^4.0.90 | provider: 'azure', model as your deployment's name, and resourceName or baseURL |
baseURLnames another endpoint. It must be https, or http only to this machine.- The adapter makes its own provider client with your key, so it ignores variables such as
ANTHROPIC_BASE_URLandOPENAI_BASE_URL. - It asks for structured output, with retries off and no tools. It asks OpenAI and Azure not to keep the request.
- A judge of your own is a module whose default export is an
EvaluatorFactory. Retest checks whatever it answers.
The check Link to The check
test.evaluate({ judge, requirement, evidence, context, mode, timeoutMs }) takes a requirement and its evidence. Await it.
requirementis one sentence, or criteria by id, such as{ saved: 'The task shows as saved.', titled: 'It shows its title.' }. Every criterion must pass.contextis reference text the judge may read, such as a policy an answer must follow.modeisrequired, the default, oradvisory.timeoutMscan shorten the check's time, never lengthen it.
| Evidence | Gives the judge |
|---|---|
{ capture: 'screenshot', app } | A screenshot of that app's page, taken as the check runs |
{ text, label } | Text the test supplies, such as a reply it read. It is redacted first. |
{ recording: { step }, app } | The frames the recording kept over the latest step with that name |
{ recording: { lastMs }, app } | The frames of the last lastMs milliseconds before the check |
{ diagnostics: 'console', app } | The console or network records of that app so far, as text |
A judge that reads frames must accept frames. Each criterion over frames names its kind. state judges the last frame. seen asks for something to appear at some point. never forbids something for the whole stretch.
await test.evaluate({ requirement: { banner: { kind: 'never', requirement: 'No error banner appears during the save.' } }, evidence: { recording: { step: 'save the task' } },})Verdicts Link to Verdicts
A check passes only when every criterion passed. A failed criterion fails it. A criterion the judge could not decide leaves it inconclusive, never a pass.
| The check ends | Required | Advisory |
|---|---|---|
| Pass | Counts as the test's assertion | Nothing more |
| Fail | The test fails, exit code 1 | A warning |
| Inconclusive | The test is inconclusive, exit code 2 | A warning |
| Error | The test is an error, exit code 2 | A warning |
- Inconclusive covers a judge that could not decide and missing evidence, such as a screenshot that failed or a recording with no frame.
- Error covers a judge that could not be set up, no answer in time, a limit reached, and an answer that breaks the contract.
- A test whose only checks are advisory makes no assertion, so it fails.
What a judge cannot do Link to What a judge cannot do
- Clear a failure. An earlier failure stays the test's failure. A later passing check clears nothing.
- Act on the app. The request holds no tool, no page and no function.
- Be overruled by the test. Catching the error changes nothing: Retest has recorded the verdict and fails the test either way.
- Pass without its evidence. A required check whose evidence is missing never passes.
The evidence travels apart from the criteria, and the judge is told that text and pixels in it are data. That lowers the risk of an app's own text steering the verdict. It does not remove it.
A screenshot check knows only what the page shows. It cannot tell whether a task was saved on the server, so check that with expect.
Cost and keys Link to Cost and keys
- By default a test makes at most 5 calls and a run at most 100, with 2 at once. A call taken is never given back, even across workers.
- Each answer is bounded to 1000 output tokens, and each request to 4 images. Nothing estimates a price.
- Keys are left out of the environment of each test's process, each app server and each browser.
- A key is written as
{{visual.apiKey}}in everything Retest records, so a provider error that quotes it leaves it nowhere in the run folder. - Retest has not yet measured how often a live judge is right.