Skip to content

Agentic Evaluations

Evaluate agents you run yourself, using their trajectories.

Run your agent in its existing runtime, upload its trial results, and let elluminate inspect and rate the recorded trajectories.

What agentic evaluation is

When to use agentic evaluations

Use this workflow when your agent makes several LLM calls, calls tools or delegates work, and you need to evaluate what it did as well as its final answer. Your runner can be your own code, Harbor, LangChain, CrewAI, AutoGen, or another framework.

For single-turn outputs, or tool calls elluminate generates itself, use the Tool Calling guide.

Core concepts

  • Task — one unit of work, identified by a unique task name.
  • Collection — the set of tasks, normally with a task column and an optional instruction column.
  • Trajectory (trace) — a step-by-step record of one run: messages, tool calls, and observations.
  • Criterion — a YES/NO question rated against a trajectory. A criterion set applies to every task; task-specific criteria apply to one task.
  • Overall Rating — the automatic YES/NO rating for overall task success.
  • Experiment — binds a collection and criterion set, then holds uploaded traces, ratings, and metrics.

The trajectory format

Your runner must produce an ATIF trajectory (Agent Trajectory Interchange Format). elluminate has no framework-specific importer: convert your runner's output to ATIF, wrap it in a trial result, and upload it through the UI or SDK.

ATIF records a run that has already happened. It is different from UCE (elluminate.uce/1), which is conversation input that elluminate runs itself; see Conversations.

UI walkthrough

Create an Agentic collection
Collections → New collection, with the Agentic collection type selected.
  1. Create an Agentic collection. Choose Agentic in Collections → New collection. Keep task as a unique Text column; it identifies every upload. instruction is optional when the trajectory already contains the input. Agentic experiments do not need a prompt template and never generate responses automatically.
  2. Create a criterion set. In Criteria Library, add clear YES/NO questions that can be answered from a trajectory. Use task-specific criteria for checks that apply to only one task.
  3. Create an Agentic experiment. In Experiments → New, choose Agentic and select the collection and criterion set.
  4. Run your agent externally and collect one ATIF trajectory per task.
  5. Upload the traces in the UI or with the SDK. With evaluation enabled, elluminate rates every criterion against each trajectory.
  6. Review the results in the UI.

Automatic Overall Rating

Every Agentic experiment adds Overall Rating to its frozen evaluation definition. It remains separate from a same-named criterion in your criterion set.

Create an Agentic experiment
Experiments → New: creating an Agentic experiment.

Import tasks from a Harbor zip

To import Harbor task folders, open an Agentic collection, choose Add task, then Upload zip. The archive may contain folders at any depth; each folder with task.toml becomes a task.

tasks/
├── fix-parser/
│   ├── task.toml
│   ├── instruction.md
│   └── criteria.toml
└── needs-doc/
    ├── task.toml
    └── instruction.md
File Required Imported as
task.toml yes Task identity and task configuration. [task].name is used when present; otherwise the folder name is used.
instruction.md yes The task instruction; it must not be empty.
criteria.toml no Criteria for that task only.
environment/, solution/, tests/ no Text-based task configuration retained for export.

The collection needs a task-name column and a Text instruction column. criteria.toml has one [[criteria]] entry per criterion. Labels are optional; when supplied, they must be unique per task, at most 25 characters, and not Overall Rating. Criterion text is limited to 4,000 characters.

Re-importing updates matching tasks instead of creating duplicates. Keep a criterion label stable to preserve its version history; renaming a label removes the old criterion and adds a new one. Tasks absent from the archive stay untouched.

Task configuration and text files under the listed directories are stored and re-exported. Binary files, files over 256 KB, and other paths are reported but not stored. Each upload is limited to 25 MB compressed, 300 tasks, 2 MB per file, 100 criteria per task, and 100 implementation files totalling 1 MB per task.

Task-specific criteria

Use the collection's Tasks view for checks that apply to one task only. Open a task's criteria sheet, add the criterion, and save. Labels are optional and unique within the task; criteria can reference task fields such as {{expected_outcome}}.

Tasks view of an Agentic collection
One card per task, including its task-specific criteria count.

An experiment evaluates its criterion set, the task's criteria, and Overall Rating. Those values are frozen when the experiment is created; create a new experiment after changing them.

Upload agent traces in the UI

Open the Agentic experiment and choose Upload agent traces. You can drop a file or paste JSON, then confirm the task-name column, choose whether to evaluate after upload, and set an epoch for another run of the same tasks.

After the file is parsed, the dialog shows which trace goes to which task, as the platform resolves it. The rest are listed under Unmatched traces, where you drag them onto a task or pick them from that task's list. A task left without a trace is skipped, and unmatched traces are not uploaded.

Upload file format

Use a .json object or array, or .jsonl with one complete object per line. Each object is an SDK AgentTrialResult:

  • task_id or task_name says which task the trace belongs to. task_id is the stronger key: it survives a rename and needs no task-name column. A trace with neither is not rejected in the UI, where you assign it after parsing.
  • trajectory contains the ATIF trajectory.
  • reward, cost_usd, steps, messages, and criterion_ratings are optional.

A trace file is not an upload file

An ATIF trace alone is not a trial result. Wrap it under trajectory.

[
  {
    "task_name": "write-hello-world",
    "reward": 1.0,
    "cost_usd": 0.0042,
    "trajectory": {
      "schema_version": "ATIF-v1.0",
      "session_id": "run-001/write-hello-world",
      "agent": {
        "name": "demo-agent",
        "version": "0.1.0",
        "model_name": "claude-sonnet-4-6"
      },
      "steps": [
        {
          "step_id": 1,
          "source": "user",
          "message": "Write a Python hello world script to hello.py"
        },
        {
          "step_id": 2,
          "source": "agent",
          "message": "Writing hello.py.",
          "tool_calls": [
            {
              "tool_call_id": "tc_1",
              "function_name": "write_file",
              "arguments": {"path": "hello.py", "content": "print('Hello, World!')"}
            }
          ],
          "observation": {
            "results": [{"source_call_id": "tc_1", "content": "wrote 22 bytes"}]
          },
          "metrics": {"prompt_tokens": 420, "completion_tokens": 61, "cost_usd": 0.0042}
        }
      ],
      "final_metrics": {
        "total_steps": 2,
        "total_cost_usd": 0.0042,
        "total_prompt_tokens": 420,
        "total_completion_tokens": 61
      }
    }
  }
]

With Evaluate after upload enabled, omit criterion_ratings and let elluminate rate the trace. For precomputed ratings, include each criterion's criterion_id from the experiment definition; labels may repeat. Older experiments can still use unambiguous labels against their live definition.

Trials are processed independently, so one invalid trajectory does not block the rest. Evaluation runs in the background.

Reviewing results in the UI

Agentic experiment overview
The experiment overview.
  • Overview shows aggregate metrics, criterion pass rates, and an experiment summary when available.
  • Sample Navigator shows a sample's phase timeline, trace, and ratings. Sub-agent traces are nested under the step that spawned them.
  • Responses Overview lists every response.
Agent Trace and Phase Timeline
Phase Timeline and Agent Trace for one sample.

SDK walkthrough

The example creates or reuses an Agentic collection, criterion set, and experiment; converts runner output to AgentTrialResult; uploads it; and waits for evaluation.

export ELLUMINATE_API_KEY=<your-key>
uv run --directory elluminate_sdk python examples/example_harbor_agentic_upload.py
"""Harbor-based Agentic Evaluation: end-to-end upload example.

This example shows the full workflow for evaluating an agent that was run
externally with Harbor (or any other agent framework), and uploading the
results, including ATIF trajectories, to elluminate for inspection and
automatic per-criterion evaluation.

Workflow:

1. Create an AGENTIC collection whose rows are the agent's tasks. The task
   name lives in a single TEXT `task` column (matched against each uploaded
   result's `task_name`); no prompt template is needed because AGENTIC
   experiments do not auto-generate responses.
2. Create a criterion set describing what counts as success.
3. Create an AGENTIC experiment (no auto-generation; results are uploaded).
4. Run the agent externally (Harbor CLI, LangChain, CrewAI, custom code).
5. Read Harbor's per-task output and convert it into `AgentTrialResult` objects.
6. Upload via `experiment.upload_agent_results(...)`.
7. The backend stores the trajectories and, when `evaluate=True` and
   trajectories are present, elluminate automatically rates each criterion
   against the trajectory.
8. Await that evaluation via `experiment.wait_for_evaluation(...)` so the
   ratings are complete before reading results back.

The script is idempotent: collections, criterion sets, and experiments are
reused across runs; uploads are skipped when an experiment already contains
responses.

For a self-contained demo this script uses a small in-memory stand-in
(`HARBOR_RUN`) for the runner's output. In a real run you would replace it
with your own loader that reads your runner's per-task output from disk — for
Harbor, the run directory at `~/.harbor/runs/<run_name>/tasks/<task>/...`.
"""

from typing import Any

from dotenv import load_dotenv
from elluminate import AgentTrialResult, Client
from elluminate.schemas import CollectionColumn, ColumnTypeEnum
from elluminate.schemas.criterion import CriterionIn
from elluminate.schemas.experiments import Experiment

load_dotenv(override=True)

client = Client()
llm_config = client.get_llm_config(name="Claude Sonnet 4.6")

# Upper bound for the evaluation wait in step 6. Size it to your run: rating is
# per trajectory, and this demo uploads two.
EVALUATION_TIMEOUT_SECONDS = 1800

# Mock "Harbor output"; in a real integration this is read from disk.
# Each entry is what a Harbor run produces per task: a short task identifier,
# the instruction text, final messages, aggregate metrics, and an ATIF
# trajectory describing every step the agent took.
HARBOR_RUN: list[dict[str, Any]] = [
    {
        "task_name": "write-hello-world",
        "instruction": "Write a Python hello world script to hello.py",
        "reward": 1.0,
        "steps": 2,
        "cost_usd": 0.0042,
        "input_tokens": 420,
        "output_tokens": 61,
        "duration_seconds": 3.2,
        "messages": [
            {"role": "user", "content": "Write a Python hello world script to hello.py"},
            {"role": "assistant", "content": "Wrote hello.py: print('Hello, World!')"},
        ],
        "trajectory": {
            "schema_version": "ATIF-v1.0",
            "session_id": "harbor-run-001/write-hello-world",
            "agent": {
                "name": "harbor-demo-agent",
                "version": "0.1.0",
                "model_name": "claude-sonnet-4-6",
            },
            "steps": [
                {
                    "step_id": 1,
                    "source": "user",
                    "message": "Write a Python hello world script to hello.py",
                },
                {
                    "step_id": 2,
                    "source": "agent",
                    "message": "Writing hello.py.",
                    "tool_calls": [
                        {
                            "tool_call_id": "tc_1",
                            "function_name": "write_file",
                            "arguments": {"path": "hello.py", "content": "print('Hello, World!')"},
                        }
                    ],
                    "observation": {
                        "results": [{"source_call_id": "tc_1", "content": "wrote 22 bytes"}],
                    },
                    "metrics": {"prompt_tokens": 420, "completion_tokens": 61, "cost_usd": 0.0042},
                },
            ],
            "final_metrics": {
                "total_steps": 2,
                "total_cost_usd": 0.0042,
                "total_prompt_tokens": 420,
                "total_completion_tokens": 61,
            },
        },
    },
    {
        "task_name": "reverse-string-function",
        "instruction": "Create a Python function that reverses a string in reverse.py",
        "reward": 0.5,
        "steps": 2,
        "cost_usd": 0.0031,
        "input_tokens": 310,
        "output_tokens": 42,
        "duration_seconds": 2.1,
        "messages": [
            {"role": "user", "content": "Create a Python function that reverses a string in reverse.py"},
            {"role": "assistant", "content": "Wrote reverse.py with a one-line slice-based reverse."},
        ],
        "trajectory": {
            "schema_version": "ATIF-v1.0",
            "session_id": "harbor-run-001/reverse-string-function",
            "agent": {
                "name": "harbor-demo-agent",
                "version": "0.1.0",
                "model_name": "claude-sonnet-4-6",
            },
            "steps": [
                {
                    "step_id": 1,
                    "source": "user",
                    "message": "Create a Python function that reverses a string in reverse.py",
                },
                {
                    "step_id": 2,
                    "source": "agent",
                    "message": "Writing reverse.py.",
                    "tool_calls": [
                        {
                            "tool_call_id": "tc_1",
                            "function_name": "write_file",
                            "arguments": {
                                "path": "reverse.py",
                                "content": "def reverse(s: str) -> str:\n    return s[::-1]\n",
                            },
                        }
                    ],
                    "observation": {
                        "results": [{"source_call_id": "tc_1", "content": "wrote 42 bytes"}],
                    },
                    "metrics": {"prompt_tokens": 310, "completion_tokens": 42, "cost_usd": 0.0031},
                },
            ],
            "final_metrics": {
                "total_steps": 2,
                "total_cost_usd": 0.0031,
                "total_prompt_tokens": 310,
                "total_completion_tokens": 42,
            },
        },
    },
]


def harbor_to_agent_trial(task_output: dict[str, Any]) -> AgentTrialResult:
    """Map one Harbor per-task output dict to an `AgentTrialResult`.

    `task_name` on `AgentTrialResult` is what elluminate matches against the
    collection's task-name column, so here we set it to the full instruction
    text (which is also what the `task` column row holds).
    """
    return AgentTrialResult(
        task_name=task_output["instruction"],
        messages=task_output["messages"],
        reward=task_output["reward"],
        steps=task_output["steps"],
        cost_usd=task_output["cost_usd"],
        input_tokens=task_output["input_tokens"],
        output_tokens=task_output["output_tokens"],
        duration_seconds=task_output["duration_seconds"],
        trajectory=task_output["trajectory"],
        metadata={"run_name": "harbor-run-001", "task_id": task_output["task_name"]},
    )


# Step 1: AGENTIC collection with a single TEXT `task` column.
# The `task` column must be TEXT so elluminate auto-selects it as the
# task-name column that uploaded results are matched against. No prompt
# template is required because AGENTIC experiments never auto-generate;
# responses are supplied by `upload_agent_results`.
collection, _ = client.get_or_create_collection(
    name="Harbor Demo Tasks",
    defaults={
        "collection_type": "AGENTIC",
        "columns": [CollectionColumn(name="task", column_type=ColumnTypeEnum.TEXT)],
        "variables": [{"task": h["instruction"]} for h in HARBOR_RUN],
    },
)

# Step 2: criterion set defining what success looks like for these tasks.
# Labels are display text and may repeat; frozen criterion IDs identify ratings.
criterion_set, _ = client.get_or_create_criterion_set(
    name="Harbor Demo Criteria",
    defaults={
        "criteria": [
            CriterionIn(
                criterion_str="Did the agent correctly complete the requested task?",
                label="task-complete",
            ),
            CriterionIn(
                criterion_str="Did the agent use tools appropriately?",
                label="uses-tools",
            ),
            CriterionIn(
                criterion_str="Is the agent's final output correct?",
                label="output-correct",
            ),
        ],
    },
)


def get_or_create_agentic_experiment(name: str, description: str) -> tuple[Experiment, bool]:
    """Return an AGENTIC experiment, creating it if missing.

    Also reports whether the experiment already has uploaded responses so the
    caller can skip a redundant upload on re-runs (avoids epoch conflicts).
    """
    try:
        experiment = client.get_experiment(name=name, fetch_responses=False)
        populated = experiment.results is not None and experiment.results.completed_epochs > 0
        return experiment, populated
    except ValueError:
        experiment = client.create_experiment(
            name=name,
            collection=collection,
            prompt_template=None,
            criterion_set=criterion_set,
            description=description,
            evaluation_mode="AGENTIC",
            llm_config=llm_config,
        )
        return experiment, False


# Step 3: AGENTIC experiment. No auto-generation; results come from Harbor.
experiment, experiment_populated = get_or_create_agentic_experiment(
    "Harbor Demo — Agent Run",
    "Harbor-run coding agent with ATIF trajectories.",
)
print(f"Experiment: {experiment.name} (id={experiment.id})")

# Step 4: Convert Harbor output to `AgentTrialResult` objects.
results = [harbor_to_agent_trial(task_output) for task_output in HARBOR_RUN]

# Step 5: upload with `evaluate=True`. elluminate rates every
# criterion against the trajectory and fills in per-criterion ratings.
if experiment_populated:
    print("Experiment already has responses; skipping upload.")
else:
    upload = experiment.upload_agent_results(
        results=results,
        evaluate=True,
    )
    print(
        f"Uploaded: {upload.created_responses} responses, "
        f"{upload.created_ratings} ratings, "
        f"{upload.pending_evaluations} pending trace evaluations"
    )
    if upload.errors:
        print(f"Errors: {upload.errors}")

    # Step 6: wait for elluminate's trace agent to finish rating the
    # uploaded trajectories, so the ratings below are complete. Rating a large
    # run takes minutes to hours; the timeout bounds the wait so a stuck
    # evaluation raises TimeoutError instead of blocking forever. Omit it to
    # wait indefinitely.
    if upload.pending_evaluations:
        print(f"Awaiting evaluation of {upload.pending_evaluations} trajectories...")
        event = experiment.wait_for_evaluation(
            task_id=upload.evaluation_task_id, timeout=EVALUATION_TIMEOUT_SECONDS
        )
        rated = event.progress.responses_rated if event.progress else 0
        print(f"Evaluation {event.status.value.lower()}: {rated} response(s) rated")

# Step 7: Verify the trajectories are queryable from the SDK.
experiment.fetch_responses()
for resp in experiment.responses():
    task = resp.prompt.template_variables.input_values.get("task", "?")
    steps = len(resp.trajectory["steps"]) if resp.trajectory else 0
    print(f"  [{task[:50]}] trajectory_steps={steps}")

Check the matching before you upload

preview_matching() answers which task each trial would be stored against, and stores nothing. It runs the matching the upload runs, so what it reports is what the upload will do.

preview = experiment.preview_agent_results_matching(results)
for item in preview.unmatched_results:
    print(item.index, item.reason)

An upload in which no result names a task of the experiment is refused whole and raises TaskMatchError, which carries the tasks it does accept:

from elluminate import TaskMatchError

try:
    experiment.upload_agent_results(results=results)
except TaskMatchError as exc:
    for task_id, task_name in ((t.task_id, t.task_name) for t in exc.accepted_tasks):
        print(task_id, task_name)

A batch where some results match is still stored, with the rest reported in errors. Upload under task_id when your runner has no stable naming scheme, or when a name is ambiguous because a row was renamed to another task's name.

Growing the dev set (the eval flywheel)

Add production failures to the collection so later runs cover them too. Adding rows expands the task set; increasing an upload epoch reruns the same tasks.

collection = client.get_collection(name="my-agent-dev-set")
collection.add_many([
    {"task": "New failing scenario discovered in prod ..."},
    {"task": "Another regression case ..."},
])

AgentTrialResult fields

Field Required Description
task_id one of the two The task's row id, from get_definition(). Authoritative when set: the result is resolved by id, with no name lookup.
task_name one of the two A task name frozen when the experiment was created, or the row's current name after a rename.
messages no Final OpenAI-format messages shown on the response page.
reward no Primary reward score (0.0–1.0).
steps no Number of agent steps or LLM calls.
cost_usd no Total USD cost; derived from the trajectory when absent.
duration_seconds no Wall-clock duration.
input_tokens no Total input tokens, including cached reads.
output_tokens no Total output tokens.
cached_tokens no Cached input tokens, already included in input_tokens.
error no Error message for a failed trial.
metadata no Free-form data shown on the response page.
trajectory no ATIF trajectory, validated by the backend.
criterion_ratings no Precomputed YES/NO ratings; use criterion_id for frozen experiments.

ATIF trajectory format

trajectory uses ATIF. The upload example above is a minimal ATIF v1 trajectory wrapped in a trial result.

How cost and tokens are resolved

elluminate uses the first available cost source: cost_usd on the trial, final_metrics.total_cost_usd, the sum of step costs (including sub-agents), then a token-based estimate from model_name. Token totals use the same precedence.

Input-token counts include cached reads. Put provider-specific cache-write tokens in metrics.extra (and run totals in final_metrics.extra); estimated costs are shown with ~. Send cost_usd or metrics.cost_usd whenever your runner knows the actual cost.