Build PAA / Evaluators
Measure performance without confusing judgment with authority
Record what an evaluator judges, how it produces the verdict, and what that verdict is allowed to do. Multiple evaluators may judge the same property without sharing technique or authority.
An unvalidated evaluator is not a gate. It is a guess with a UI.
Evaluator fields
A score alone cannot say what was judged or what decisions may rely on it.
| Field | Question |
|---|---|
property | The named property this evaluator produces a verdict about. |
target | The layer inspected: input, process, output, or outcome. |
technique | What produces the verdict: deterministic, classifier, llm_judge, or human. |
evaluation_basis | The criteria, reference, or procedure that supports the verdict: invariant, reference_label, rubric, human_gold, or downstream_result. |
epistemic_status | Whether the verdict is proxy or ground_truth. |
version | The identity of this evaluator instance. Comparable evaluators may assert the same property without sharing technique, evaluation basis, epistemic status, version, or authority. |
authority | Whether this evaluator's verdict can halt the governed effect: advisory or blocking. Placement, when it runs, comes from the task's position_policy, not from the evaluator. |
Target
Evaluate at the cheapest layer that actually catches the failure.
| Layer | Question | Cost | Note |
|---|---|---|---|
| Input | Was the right context present and well-formed before acting? | Cheapest | Catches structural failures before work starts. |
| Process | Were the present inputs used correctly during generation? | Hard / often sealed | Usually inferred from output and outcome. |
| Output | Is the produced artifact good against some standard? | Needs an evaluation basis | The default eval layer. |
| Outcome | Did it work downstream? | Truest, most lagged | Noisy and delayed, but the real numerator. |
Technique
Four closed values. Ordered by cost, cheapest first.
| Technique | Examples | Cost / latency | Best use |
|---|---|---|---|
deterministic | Schemas, rules, bounds, assertions, allow/deny lists, reference validators | Zero | Structural correctness and hard invariants |
classifier | Learned models on embeddings, fine-tuned classifiers, distilled decision models | Cheap inference | The mature high-volume evaluator after sufficient label accumulation |
llm_judge | A model scores output against a rubric, with or without a reference | Expensive / latent | Bootstrap labels before distilling to cheaper evaluators |
human | Expert review, authorization decisions, quality labels | Highest | Ground truth for ambiguous or high-stakes cases; authorization gates |
Evaluation basis
The evaluation basis is the criteria, reference, or procedure that supports the
evaluator's verdict. It is a nested field, { kind, ref }, not a
single closed word. Five closed kind values:
| Kind | What it means |
|---|---|
invariant | A structural rule or assertion that must hold. |
reference_label | A comparison against a known reference label. |
rubric | Written criteria a technique applies to produce a verdict. A rubric-driven LLM judge is commonly proxy, rubric is the basis, proxy is the separately-declared epistemic status. |
human_gold | Human expert judgment used as the reference. Human gold ordinarily carries ground_truth status, but the two fields vary independently. |
downstream_result | A measured downstream outcome. |
Epistemic status
Epistemic status is a separate field from evaluation basis: it declares whether the verdict is an approximation of ground truth or ground truth itself. Two closed values:
proxy: The verdict is an approximation of ground truth.ground_truth: The verdict is treated as ground truth itself.
An LLM judge can use a rubric as its basis while remaining a proxy. The rubric
says what it compares against; proxy says how much weight the verdict carries.
Authority
Authority is the gate effect of an evaluator: whether its verdict can halt the governed effect. Two closed values:
advisory: The verdict is recorded and surfaced but cannot halt the governed effect.blocking: When the task placement is blocking, the verdict can halt the current governed effect and all blocking evaluators must pass. With async or offline placement, the verdict can affect a later cycle or demotion instead.
A human reviewer can be advisory after execution; a deterministic check can block before
execution. shadow is a task deployment mode, not evaluator authority.
Multiple evaluators, one property
A property is not limited to one evaluator. response_quality commonly
appears twice on the same declaration: one version is an llm_judge against a
rubric, epistemic status proxy, authority advisory; a second
version is human against human_gold, epistemic status
ground_truth, a governance designation rather than a claim of infallibility;
authority is also advisory while the judge is being
validated. Both evaluators run and both are recorded. Comparable evaluators may assert
the same property without sharing technique, evaluation basis, epistemic status, version,
or authority.
Evaluator succession changes which version holds blocking authority, and it does so only after measured agreement between the candidate and the evaluator it would replace; it does not mutate historical evidence. Every evidence record retains the complete identity of the evaluator that produced it, so a later authority handoff never rewrites what an earlier version actually verified.
Placement (task level)
position_policy is a top-level task field, not an evaluator field. It maps
the task's current autonomy position to a placement value: blocking, async, offline.
The policy is fixed in the declaration; the task's current position changes at runtime.
Placement describes where an evaluator participates in task execution. The task declares it through position_policy.
At HOTL, a human evaluator can leave the synchronous path while machine checks keep blocking. Per-evaluator placement overrides express that task policy. See Schema: Position policy for the selector form and its validation rules.
| Autonomy position | Typical placement value | Timing | Effect |
|---|---|---|---|
manual | offline | Batch / aggregate | Periodic evaluation; no per-execution gate. The human still performs the governed effect. |
hitl | blocking | Before execution | Halts the write until all blocking evaluators clear. |
hotl | async | After execution | Watches outputs and triggers demotion or alert. |
autonomous | offline | Batch / aggregate | Measures whether the task has earned more autonomy. |
Economics of evaluation
Evaluation is one component of the configuration's operating cost.
A cheap evaluator that misses important failures can increase total cost. An expensive evaluator may be justified when it reduces human review, failures, or remediation. There is no universal requirement that evaluation cost less than worker inference.
Operating cost = worker + evaluation + human review + retries/escalation/remediation.
Effective cost = operating cost / accepted outcomes.
These are useful accounting models, not mandatory PAA schema fields. Declare acceptance before measuring, and use the same configuration, population, and window on both sides of the ratio. Include failed and retried attempts once. With no accepted outcomes, effective cost is undefined; report spend and the zero count.
State which components are measured, estimated, or unavailable. Missing costs are not zero. If only inference is measured, call the metric inference cost per accepted outcome. See the illustrative comparison.
| Principle | Rule | Implication |
|---|---|---|
| Evaluator economics | Justify evaluation against task value, risk, and total operating cost. | An evaluator may cost more than worker inference and still reduce review, failure, or remediation costs. |
| Cost curve | Expensive at cold start, cheap at volume. | A reviewer-heavy phase can fund itself because the volume is still low. |
| Distillation | Move from llm_judge to classifier once labels exist. | The accumulated review stream pays for the cheaper evaluator you will need later. |
| Boundary discipline | Instrument from the start. | The data must already be there when the evaluator needs to get cheaper. Instrumentation is the entry criterion for manual, the first on-spectrum position, not something added later. |
Maturity curve
People and flexible judges do most early evaluation. Their verdicts create the labels needed to train a cheaper evaluator. Evaluator maturity does not automatically change a task's autonomy position.
| Stage | Evaluator setup | Evaluation role | Meaning |
|---|---|---|---|
| Cold start | human plus llm_judge; epistemic status is ground_truth or proxy pending validation | Advisory and heavily checked against human judgment | Not sufficient alone for authority-changing decisions. |
| Growing evidence | Distilled classifier begins to replace the judge; evaluation basis shifts toward reference_label or rubric signals | May support defined governance decisions after demonstrating agreement and stability | Review labels become validation and training data while authority remains explicit. |
| Mature volume | Cheap classifier on the hot path; proxy epistemic status validated against downstream outcomes | Trusted for a broader evaluation role within declared authority bounds | May reduce evaluation cost and human-review burden without granting worker authority. |
| Regression | Fallback to llm_judge or human after agreement or calibration drifts past the bar | Narrowed or suspended pending renewed validation | Evaluator authority and review burden change without automatically changing task autonomy. |
Metric choice
Optimize according to failure cost, not preference.
| Failure profile | Metric to favor | Why |
|---|---|---|
| Rare, high-cost failures | Recall | Misses are catastrophic, so escalate liberally. |
| Common, low-cost failures | Precision and throughput | False escalations become the expensive error. |
| Bootstrap labels | Human agreement | Use the judge to create the training set, not to pretend the evaluator is already proven. |
How an evaluator gets promoted and demoted
Compare a candidate evaluator with a defined reference. Once agreement and calibration clear the bar, a cheaper technique can take on a declared role; drift narrows that role or returns work to human review. Historical evidence keeps the identity that produced it. See Evidence and labels and Improvement for the two consumers of those verdicts.