Documentation / guides
Pluto qualification tooling
Roll Evals runs into capability-aware tables, scorecards, and profile dispositions with Pluto.
Pluto is the code-backed evaluation tooling that composes Evals into a larger
qualification result. It does not replace eval.Run: each runnable table
expands to one eval.Suite, uses the table’s evaluator set, and retains the raw
per-table eval.Report behind the rollup. The underlying lifecycle is the Evals
run and result contract.
Pluto adds capability-aware qualification
The core objects are:
| Object | Contract |
|---|---|
qual.Manifest | Secret-free target identity, provider/model metadata, endpoint class, revision, and capabilities. |
qual.Table | A named, versioned scenario family with one dimension, required capabilities, and evaluators. |
qual.Pack | A versioned set of tables with unique table names and scenario IDs. |
qual.TablePlan | Preflight output that either carries a runnable suite and evaluators or records missing capabilities. |
qual.Scorecard | Manifest plus executed and skipped TableResult values. |
Manifest fingerprints are stable hashes over canonical manifest JSON. Credentials are not part of the manifest. Endpoint class limits what a result can claim to have observed: remote targets expose requests, responses, tools, usage, timing, and errors, while process targets can add process-level behavior through their own target adapter.
Plan before execution
qual.Plan(pack, manifest) validates both inputs, checks each table’s required
capabilities, and preserves a non-runnable plan with its missing capabilities.
It performs no target work. A capability skip is visible coverage, not a dropped
case.
plans, err := qual.Plan(pack, manifest)
if err != nil {
return err
}
for _, plan := range plans {
if !plan.Runnable {
log.Printf("skip %s: missing %v", plan.Table, plan.Missing)
continue
}
// Table.Suite expands exactly this table into the eval.Run input.
report, err := eval.Run(ctx, eval.RunConfig{}, plan.Suite, target, plan.Evaluators...)
if err != nil {
return err
}
_ = report
}
The shared pkg/run.Execute path performs this plan-to-run transition for both
offline fixture targets and live per-table targets. Its Spec requires exactly
one of a shared Target or a TargetForTable factory, so a caller cannot
accidentally provide both or neither.
Scorecard keeps coverage visible
Scorecard.Dimensions() rolls results up by dimension in stable name order:
passcontributes one verdict and one pass.failcontributes one verdict and no pass.unverifiedanderrorcontribute to the assessment denominator but not to quality or verdict count.skippedtable results increaseSkippedTablesand do not enter the assessment denominator.- A dimension with no verdicts is
Undecided, never a silent zero and never a silent pass.
The score is 100 * passes / verdicts. Coverage is verdicts / assessments,
so a high score with weak evidence remains visible as low coverage. The raw
reports remain available for finding counts, severity counts, and case-level
inspection.
dims, err := result.Scorecard.Dimensions()
if err != nil {
return err
}
for _, dim := range dims {
fmt.Printf("%s score=%.1f coverage=%.2f undecided=%t\n",
dim.Dimension, dim.Score, dim.Coverage, dim.Undecided)
}
Profiles derive dispositions
profile.Profile contains mandatory requirements and optional restrictions.
profile.Evaluate is pure policy over a scorecard. Its precedence is explicit:
- Any violated mandatory requirement yields
Rejected. - Otherwise, any undecided mandatory requirement yields
Unverified. - Otherwise, an unmet restriction yields
Restricted. - Otherwise, the result is
Qualified.
A minimum coverage requirement that cannot be met because evidence is missing is undecided, not a demonstrated quality violation. A finding or severity count can be a separate requirement with a maximum bound.
minScore := 80.0
profile := profile.Profile{
Name: "production-agent", Revision: "v1",
Requirements: []profile.Requirement{{
Dimension: "capability", MinScore: &minScore,
}},
}
result, err := profile.Evaluate(card, profile)
if err != nil {
return err
}
if result.Disposition != profile.Qualified {
return fmt.Errorf("qualification disposition: %s", result.Disposition)
}
Disposition.Rank provides a separate worst-to-best ordering for callers that
need a minimum acceptable floor. It does not alter Evaluate’s derivation
precedence.
Run through the shared execution core
Use run.Execute when you want Pluto to retain skipped plans, reports, and
partial results in one Result:
result, err := run.Execute(ctx, run.Spec{
Manifest: manifest,
Packs: []qual.Pack{pack},
Target: target, // A scripted target works for deterministic tests.
Config: eval.RunConfig{Trials: 2},
})
if err != nil {
// A non-nil error means the overall execution did not complete, but result
// may still carry reports for tables that already finished.
return err
}
scorecard := result.Scorecard
For live tables, run.BuildTarget combines a caller-supplied inference client,
the table environment template, the manifest model, and the table revision into
the Evals Inference target. For
Go tests, plutotest.Run wraps the same execution core and
plutotest.RequireDisposition gates an allowed set of profile outcomes.
Persist a Pluto result
pluto/pkg/reportjson emits pluto-report/v1. It records the manifest and
fingerprint, dimension scores and coverage, status rollup, each table’s
skipped/runnable state, optional profile result, and the embedded bytes from
Evals’ own redacted report/v1 codec. Decoding a table report therefore keeps
the same redaction behavior described in Evals reporting.
Source
Planning and table expansion are implemented in pkg/qual/pack.go. Shared execution is in pkg/run/run.go, scorecard rollups are in pkg/qual/scorecard.go, and profile derivation is in pkg/profile/evaluate.go.
Proof
Planning, capability skips, and suite expansion are covered by pkg/qual/pack_test.go. Score and coverage semantics are covered by pkg/qual/scorecard_test.go, and disposition precedence by pkg/profile/evaluate_test.go. The shared run and Go test wrappers are pkg/run/run.go and pkg/plutotest/run.go.