evaluate() runs a module on every row of a data frame (as
run_dataset() does), scores each output with a metric, and returns the
mean score together with the per-row scores, predictions and errors.
Arguments
- module
A module, such as one created with
module()ormodule_fn().- ...
Arguments for the evaluation:
data(required, second argument): a data frame with one column per signature input, plus the columns the metric compares against.metric(required): a function called asmetric(prediction, expected)for each row, wherepredictionis the row's output (a named list, as returned byrun()) andexpectedis the whole data row as a one-row data frame. It returns a logical or numeric score, orlist(score = , feedback = ). Built-in metrics takefieldto name the output and column to compare, as inmetric_exact_match(field = "sentiment"). Metrics made withmetric_with_trace()also receive the row's execution trace.epochs: how many times to run every row (default1L). Epochs after the first use their own cache partition, so each samples fresh responses. Repeating the sameevaluate()call replays all epochs from the cache; pass.cache = FALSEto sample again..cache:NULL(the default) followsconfigure_cache();FALSEskips the response cache..llm,.concurrency,.progress,.trace_context: as inrun_dataset()..return_format:"structured"(the default) or"simple", which leaves outmetadata,traces,epoch_tracesanddata.
Value
A list of class dsprrr_evaluation with:
mean_score: the mean score, with failed rows counted as 0. Withepochs > 1, the mean over every row and epoch.scores: one score per row (logical scores become 0 or 1),NAfor failed rows. Withepochs > 1, each row's mean across epochs,NAif any epoch failed.predictions: one output per row.n_evaluated: rows with a score.n_errors: rows where the call or the metric failed, with the messages inerrors.n_run_errors/run_errorsandn_metric_errors/metric_errorsseparate the two kinds.total_cost: total cost of the model calls,NAwhen any cost is unknown.feedbacks: the feedback text from metrics that returnlist(score = , feedback = )(seemetric_with_feedback()), otherwiseNA.program_artifact_id,trace_context: the program's identity and the caller's correlation context.metadata: one list of call metadata per row.traces: one execution trace per row, as passed to trace-aware metrics. Traces can contain prompts, inputs and responses.data: the result ofrun_dataset(.return_format = "structured").
With epochs > 1, the list also has epoch_scores (one score vector per
epoch), score_std (the standard deviation of the epoch means), ci_95
(a 95% confidence interval for mean_score) and epoch_traces; the
predictions, metadata, traces and data come from the last epoch.
An empty data gives a warning and an NA mean score.
Details
A row whose call fails, or whose metric errors, gets an NA score and is
counted in n_errors; mean_score counts it as 0, so failures lower the
score instead of disappearing from it. Metric errors also raise a warning.
See also
optimize_grid() and compile(), which use evaluate() to
compare candidate programs.
Other execution:
concurrency_control(),
predict.Module(),
run(),
run_async(),
run_dataset(),
run_stream(),
stream_async(),
stream_listener()
Other metrics:
as_dsprrr_metric(),
metric_contains(),
metric_custom(),
metric_exact_match(),
metric_f1(),
metric_field_match(),
metric_threshold(),
metric_with_feedback(),
metric_with_trace(),
vitals_metrics
Examples
# A keyword rule stands in for a model, so this example runs offline
rule <- module_fn(
"text -> sentiment",
function(text) {
if (grepl("love|great", text, ignore.case = TRUE)) "positive" else "negative"
}
)
testset <- data.frame(
text = c("I love it!", "Awful.", "Great value", "Not great"),
sentiment = c("positive", "negative", "positive", "negative")
)
result <- evaluate(rule, testset, metric = metric_exact_match(field = "sentiment"))
result
#>
#> ── DSPrrr Evaluation Results
#> ✔ Mean Score: 0.75
#> Evaluated: 4
#> Scores: 1, 1, 1, 0
result$scores
#> [1] 1 1 1 0
# A custom metric gets the prediction and the whole data row
same_label <- function(prediction, expected) {
prediction$sentiment == expected$sentiment
}
evaluate(rule, testset, metric = same_label)$mean_score
#> [1] 0.75
if (FALSE) { # \dontrun{
classifier <- module(
signature("text -> sentiment: enum('positive', 'negative')")
)
llm <- ellmer::chat_openai(model = "gpt-6-luna")
result <- evaluate(
classifier,
testset,
metric = metric_exact_match(field = "sentiment"),
.llm = llm
)
result$n_errors
# Run every row three times to see how much the score varies
repeated <- evaluate(
classifier,
testset,
metric = metric_exact_match(field = "sentiment"),
.llm = llm,
epochs = 3L
)
repeated$ci_95
} # }