metric_with_feedback() marks a metric whose function returns textual
feedback along with its score, as
list(score = , feedback = "what went wrong"). Feedback-aware optimizers
such as GEPA() use the feedback to guide their reflection step, as in
DSPy's GEPA. Everywhere else, including evaluate(), only the score is
used; evaluate() also returns the feedback in feedbacks.
Arguments
- fn
A function called as
fn(prediction, expected), wherepredictionis the output (a named list) andexpectedthe whole data row, as a one-row data frame. It returns a logical or numeric score, orlist(score = , feedback = )withfeedbacka single string.- field
The name of the data column that holds the expected output. It is only stored, in the metric's
"field"attribute, for optimizers that look it up;fnstill receives the whole row and prediction.
Examples
graded <- metric_with_feedback(
function(prediction, expected) {
if (identical(prediction$answer, expected$answer)) {
list(score = 1, feedback = "Correct.")
} else {
list(
score = 0,
feedback = paste0("Expected '", expected$answer, "' but got '", prediction$answer, "'.")
)
}
},
field = "answer"
)
row <- data.frame(question = "What is 2 + 2?", answer = "4")
graded(list(answer = "4"), row)
#> $score
#> [1] 1
#>
#> $feedback
#> [1] "Correct."
#>
graded(list(answer = "5"), row)
#> $score
#> [1] 0
#>
#> $feedback
#> [1] "Expected '4' but got '5'."
#>