Skip to documentation
Documentation navigation

Documentation navigation

Documentation / guides

Evals

Build trustworthy qualification runs from cases, targets, evaluators, and reports.

developer

Evals turns an interaction into an auditable qualification result. A Scenario describes one case, a Target produces an Observation, an Evaluator turns that observation into an Assessment, and Run preserves every result in a Report. The packages under eval/exact, eval/judge, and eval/reportjson add concrete evaluators, model judges, and a redacted report sink without changing that core contract.

The useful mental model is a data pipeline, not a callback that returns a boolean:

%%{init: {"theme":"dark"}}%%
sequenceDiagram
    participant S as Suite
    participant R as eval.Run
    participant T as Target
    participant O as Observation
    participant E as Evaluators
    participant P as Report
    S->>R: scenarios + RunConfig
    R->>T: Observe(ctx, Scenario)
    T-->>O: conversation + trace evidence
    R->>E: Evaluate(ctx, Sample)
    E-->>R: Assessment or stage error
    R-->>P: ordered SampleReports + summary

Use Cases and runs for the case-to-report lifecycle, Evaluator contracts for quality gates, Reporting for durable output, and Integration when an application supplies a model, agent, or qualification rollup.

Choose the layer you need

NeedStart hereContract to keep in mind
Describe cases and expected behaviorCases and suitesIDs, revisions, typed input, and optional expectations are validated before work starts.
Execute cases repeatedly or concurrentlyRuns and resultsResults stay in scenario-major, trial-minor order. Cancellation returns completed work with its context error.
Write a gate or quality checkEvaluator contractsA quality miss is fail; evaluator infrastructure trouble is error; missing evidence is unverified.
Match or forbid exact outputExact evaluatorsText checks inspect assistant text blocks only, while tool checks inspect typed tool-use blocks.
Score ambiguous behaviorModel judgesThe judge asks for strict structured output and validates score and quote provenance locally.
Persist a safe resultReportingreport/v1 is canonical and redacted; FileSink makes the final rename atomic.
Feed a live model or agentComposition and testingSupply an eval.Target at the composition boundary, then keep deterministic fixtures beside opt-in live tests.
Roll up qualification dimensionsPluto toolingPluto composes eval.Run into capability-aware tables, scorecards, and profile dispositions.

For model identity and request construction, pair this guide with Inference’s model selection reference. When a request uses a schema, the structured output reference explains the provider-facing feature that the Evals target records as evidence.

A first complete run

The target below is deliberately small. Production targets can wrap an agent, an HTTP service, a process, or the public target/inference.NewTarget adapter. The evaluator sees only a validated Sample, so the same gate can run against a deterministic fixture and a live provider.

package evalquickstart_test

import (
	"context"
	"testing"

	"github.com/looprig/core/content"
	"github.com/looprig/eval"
	"github.com/looprig/eval/exact"
)

type fixedTarget struct{}

func (fixedTarget) Name() string { return "capital-answer" }

func (fixedTarget) Observe(_ context.Context, sc eval.Scenario) (eval.Observation, error) {
	// Keep the caller's scenario unchanged and append the target's reply to a
	// fresh conversation slice.
	conversation := append(content.AgenticMessages(nil), sc.Input...)
	conversation = append(conversation, &content.AIMessage{Message: content.Message{
		Role: content.RoleAssistant,
		Blocks: []content.Block{&content.TextBlock{Text: "Paris is the capital of France."}},
	}})
	return eval.Observation{
		Conversation: conversation,
		Scope:        eval.ScopeCase,
		Subject: eval.Subject{
			ID: "fixed-target", Kind: eval.SubjectAgent,
			Name: sc.Name, Revision: sc.Revision,
		},
	}, nil
}

func TestQualification(t *testing.T) {
	suite := eval.Suite{
		Name: "capital-smoke", Revision: "v1",
		Scenarios: []eval.Scenario{{
			ID: "france-capital", Name: "capital-answer", Revision: "v1",
			Input: content.AgenticMessages{&content.UserMessage{Message: content.Message{
				Role: content.RoleUser,
				Blocks: []content.Block{&content.TextBlock{Text: "What is the capital of France?"}},
			}}},
		}},
	}

	report, err := eval.Run(context.Background(), eval.RunConfig{}, suite, fixedTarget{},
		exact.RequiredText("Paris"), exact.ForbiddenText("London"))
	if err != nil {
		t.Fatal(err)
	}
	if report.Summary.Assessments[eval.StatusPass] != 2 {
		t.Fatalf("pass assessments = %d, want 2", report.Summary.Assessments[eval.StatusPass])
	}
}

The example has two independent assessments. A report can therefore say that the answer included the required fact while also recording a separate safety or operational failure. That separation is why downstream gates should inspect assessment status and evidence instead of treating a single scalar as the whole truth.

Source

The public pipeline is declared in run.go, with case identity in scenario.go and the exact example in examples/exact/example_test.go.

Proof

The pipeline and its invariants are implemented in run.go, scenario.go, observation.go, assessment.go, and report.go. The deterministic exact gate is exercised by examples/exact/example_test.go.

← back to documentation