Skip to contents

metric_with_feedback() marks a metric whose function returns textual feedback along with its score, as list(score = , feedback = "what went wrong"). Feedback-aware optimizers such as GEPA() use the feedback to guide their reflection step, as in DSPy's GEPA. Everywhere else, including evaluate(), only the score is used; evaluate() also returns the feedback in feedbacks.

Usage

metric_with_feedback(fn, field = NULL)

Arguments

fn

A function called as fn(prediction, expected), where prediction is the output (a named list) and expected the whole data row, as a one-row data frame. It returns a logical or numeric score, or list(score = , feedback = ) with feedback a single string.

field

The name of the data column that holds the expected output. It is only stored, in the metric's "field" attribute, for optimizers that look it up; fn still receives the whole row and prediction.

Value

A metric function of class dsprrr_feedback_metric.

Examples

graded <- metric_with_feedback(
  function(prediction, expected) {
    if (identical(prediction$answer, expected$answer)) {
      list(score = 1, feedback = "Correct.")
    } else {
      list(
        score = 0,
        feedback = paste0("Expected '", expected$answer, "' but got '", prediction$answer, "'.")
      )
    }
  },
  field = "answer"
)
row <- data.frame(question = "What is 2 + 2?", answer = "4")
graded(list(answer = "4"), row)
#> $score
#> [1] 1
#> 
#> $feedback
#> [1] "Correct."
#> 
graded(list(answer = "5"), row)
#> $score
#> [1] 0
#> 
#> $feedback
#> [1] "Expected '4' but got '5'."
#>