AI engine

AI evaluator

The judge that scores each synthetic conversation. Text only — it reads transcripts, never screenshots. A run snapshots this selection at launch, so a change affects the next run and never re-scores a finished one.

Router

No models loaded. This router needs a key to list them — enter it above and the list reloads when you leave the field, or enter the model id by hand below.

Judge model scores transcripts against expected behaviours