Skip to content
Redline
Esc
navigateopen⌘Jpreview
On this page

The vocabulary

Ten words, used exactly. Using them loosely makes a result unreadable.

Redline uses a small number of words in a precise way, and the whole product is built from them. Learning them once saves guessing later.

Word What it means here
Agent The thing being tested. Either from the catalog — a coding CLI Redline installs and drives inside a container — or your own, written with our SDK and connected by redline dev running in your process.
Task The instructions an agent is given, plus any files that belong to those instructions. Stored once, usable in any number of experiments.
Rubric A list of criteria a run is judged against. Each criterion gets its own pass or fail and its own note.
Asset Something an experiment hands to every agent in it: an MCP server, a skill, a CLI, an SDK package, or a plugin.
File A repository cloned into the machine’s workspace, or an uploaded file written into it.
Machine The computer a run happens on — a container of a chosen shape, created for the run and destroyed after it.
Experiment A cross product: every selected task, attempted by every selected agent, with those attachments, on that machine shape. Frozen once launched.
Run One task × one agent. It has a transcript, a final answer, and a verdict.
Trial An allowed retry of a failing run — not a copy of a passing one.
Schedule A saved experiment configuration re-launched on a cron. What it produces is an ordinary experiment.

The relationships worth knowing

A task does not own its machine or its rubric

Both are chosen when you launch. That is deliberate: it is what lets the same instructions be run on a bigger box, or judged by a stricter standard, without rewriting them.

An asset belongs to the experiment, not to an agent

Attaching an MCP server means every agent in that experiment gets it. Otherwise you are not comparing agents, you are comparing setups.

A file can belong to either

Attach it to the task when it is part of what the task means (“fix the bug in this repository”), or to the experiment when it is part of this particular launch. A run receives both sets, merged and de-duplicated.

A rubric is copied onto every run it judges

Each run keeps its own copy of the criteria it was actually judged against, so editing the rubric later never quietly rewrites a result you already read.

Why trials are not copies

Setting trials to three means try up to three times. A task that passes on the first attempt produces one run, not three identical passes.

That is what makes an unreliable agent visible. Passing on the third attempt is a different result from passing on the first, and both are in the record — the run row shows 3/3 attempts and the transcript contains every one of them.

Was this page helpful?