Build PAA / Evaluators

Measure performance without confusing judgment with authority

Record what an evaluator judges, how it produces the verdict, and what that verdict is allowed to do. Multiple evaluators may judge the same property without sharing technique or authority.

An unvalidated evaluator is not a gate. It is a guess with a UI.

Evaluator fields

A score alone cannot say what was judged or what decisions may rely on it.

Evaluator identity fields
FieldQuestion
propertyThe named property this evaluator produces a verdict about.
targetThe layer inspected: input, process, output, or outcome.
techniqueWhat produces the verdict: deterministic, classifier, llm_judge, or human.
evaluation_basisThe criteria, reference, or procedure that supports the verdict: invariant, reference_label, rubric, human_gold, or downstream_result.
epistemic_statusWhether the verdict is proxy or ground_truth.
versionThe identity of this evaluator instance. Comparable evaluators may assert the same property without sharing technique, evaluation basis, epistemic status, version, or authority.
authorityWhether this evaluator's verdict can halt the governed effect: advisory or blocking. Placement, when it runs, comes from the task's position_policy, not from the evaluator.

Target

Evaluate at the cheapest layer that actually catches the failure.

Evaluator target layers
Layer Question Cost Note
Input Was the right context present and well-formed before acting? Cheapest Catches structural failures before work starts.
Process Were the present inputs used correctly during generation? Hard / often sealed Usually inferred from output and outcome.
Output Is the produced artifact good against some standard? Needs an evaluation basis The default eval layer.
Outcome Did it work downstream? Truest, most lagged Noisy and delayed, but the real numerator.

Technique

Four closed values. Ordered by cost, cheapest first.

Evaluator technique catalog
Technique Examples Cost / latency Best use
deterministic Schemas, rules, bounds, assertions, allow/deny lists, reference validators Zero Structural correctness and hard invariants
classifier Learned models on embeddings, fine-tuned classifiers, distilled decision models Cheap inference The mature high-volume evaluator after sufficient label accumulation
llm_judge A model scores output against a rubric, with or without a reference Expensive / latent Bootstrap labels before distilling to cheaper evaluators
human Expert review, authorization decisions, quality labels Highest Ground truth for ambiguous or high-stakes cases; authorization gates

Evaluation basis

The evaluation basis is the criteria, reference, or procedure that supports the evaluator's verdict. It is a nested field, { kind, ref }, not a single closed word. Five closed kind values:

Evaluation basis kinds
Kind What it means
invariant A structural rule or assertion that must hold.
reference_label A comparison against a known reference label.
rubric Written criteria a technique applies to produce a verdict. A rubric-driven LLM judge is commonly proxy, rubric is the basis, proxy is the separately-declared epistemic status.
human_gold Human expert judgment used as the reference. Human gold ordinarily carries ground_truth status, but the two fields vary independently.
downstream_result A measured downstream outcome.

Epistemic status

Epistemic status is a separate field from evaluation basis: it declares whether the verdict is an approximation of ground truth or ground truth itself. Two closed values:

  • proxy: The verdict is an approximation of ground truth.
  • ground_truth: The verdict is treated as ground truth itself.

An LLM judge can use a rubric as its basis while remaining a proxy. The rubric says what it compares against; proxy says how much weight the verdict carries.

Authority

Authority is the gate effect of an evaluator: whether its verdict can halt the governed effect. Two closed values:

  • advisory: The verdict is recorded and surfaced but cannot halt the governed effect.
  • blocking: When the task placement is blocking, the verdict can halt the current governed effect and all blocking evaluators must pass. With async or offline placement, the verdict can affect a later cycle or demotion instead.

A human reviewer can be advisory after execution; a deterministic check can block before execution. shadow is a task deployment mode, not evaluator authority.

Multiple evaluators, one property

A property is not limited to one evaluator. response_quality commonly appears twice on the same declaration: one version is an llm_judge against a rubric, epistemic status proxy, authority advisory; a second version is human against human_gold, epistemic status ground_truth, a governance designation rather than a claim of infallibility; authority is also advisory while the judge is being validated. Both evaluators run and both are recorded. Comparable evaluators may assert the same property without sharing technique, evaluation basis, epistemic status, version, or authority.

Evaluator succession changes which version holds blocking authority, and it does so only after measured agreement between the candidate and the evaluator it would replace; it does not mutate historical evidence. Every evidence record retains the complete identity of the evaluator that produced it, so a later authority handoff never rewrites what an earlier version actually verified.

Placement (task level)

position_policy is a top-level task field, not an evaluator field. It maps the task's current autonomy position to a placement value: blocking, async, offline. The policy is fixed in the declaration; the task's current position changes at runtime.

Placement describes where an evaluator participates in task execution. The task declares it through position_policy.

At HOTL, a human evaluator can leave the synchronous path while machine checks keep blocking. Per-evaluator placement overrides express that task policy. See Schema: Position policy for the selector form and its validation rules.

Position policy values by autonomy position
Autonomy position Typical placement value Timing Effect
manual offline Batch / aggregate Periodic evaluation; no per-execution gate. The human still performs the governed effect.
hitl blocking Before execution Halts the write until all blocking evaluators clear.
hotl async After execution Watches outputs and triggers demotion or alert.
autonomous offline Batch / aggregate Measures whether the task has earned more autonomy.

Economics of evaluation

Evaluation is one component of the configuration's operating cost.

A cheap evaluator that misses important failures can increase total cost. An expensive evaluator may be justified when it reduces human review, failures, or remediation. There is no universal requirement that evaluation cost less than worker inference.

Operating cost = worker + evaluation + human review + retries/escalation/remediation.

Effective cost = operating cost / accepted outcomes.

These are useful accounting models, not mandatory PAA schema fields. Declare acceptance before measuring, and use the same configuration, population, and window on both sides of the ratio. Include failed and retried attempts once. With no accepted outcomes, effective cost is undefined; report spend and the zero count.

State which components are measured, estimated, or unavailable. Missing costs are not zero. If only inference is measured, call the metric inference cost per accepted outcome. See the illustrative comparison.

Evaluator economics
Principle Rule Implication
Evaluator economics Justify evaluation against task value, risk, and total operating cost. An evaluator may cost more than worker inference and still reduce review, failure, or remediation costs.
Cost curve Expensive at cold start, cheap at volume. A reviewer-heavy phase can fund itself because the volume is still low.
Distillation Move from llm_judge to classifier once labels exist. The accumulated review stream pays for the cheaper evaluator you will need later.
Boundary discipline Instrument from the start. The data must already be there when the evaluator needs to get cheaper. Instrumentation is the entry criterion for manual, the first on-spectrum position, not something added later.

Maturity curve

People and flexible judges do most early evaluation. Their verdicts create the labels needed to train a cheaper evaluator. Evaluator maturity does not automatically change a task's autonomy position.

Evaluator maturity curve
Stage Evaluator setup Evaluation role Meaning
Cold start human plus llm_judge; epistemic status is ground_truth or proxy pending validation Advisory and heavily checked against human judgment Not sufficient alone for authority-changing decisions.
Growing evidence Distilled classifier begins to replace the judge; evaluation basis shifts toward reference_label or rubric signals May support defined governance decisions after demonstrating agreement and stability Review labels become validation and training data while authority remains explicit.
Mature volume Cheap classifier on the hot path; proxy epistemic status validated against downstream outcomes Trusted for a broader evaluation role within declared authority bounds May reduce evaluation cost and human-review burden without granting worker authority.
Regression Fallback to llm_judge or human after agreement or calibration drifts past the bar Narrowed or suspended pending renewed validation Evaluator authority and review burden change without automatically changing task autonomy.

Metric choice

Optimize according to failure cost, not preference.

Metric choice by failure profile
Failure profile Metric to favor Why
Rare, high-cost failures Recall Misses are catastrophic, so escalate liberally.
Common, low-cost failures Precision and throughput False escalations become the expensive error.
Bootstrap labels Human agreement Use the judge to create the training set, not to pretend the evaluator is already proven.

How an evaluator gets promoted and demoted

Compare a candidate evaluator with a defined reference. Once agreement and calibration clear the bar, a cheaper technique can take on a declared role; drift narrows that role or returns work to human review. Historical evidence keeps the identity that produced it. See Evidence and labels and Improvement for the two consumers of those verdicts.

Read more