Task declaration
The versioned contract for the boundary, evaluators, deployment, position policy, promotion, and demotion.
Build PAA / Framework
A task declaration is not the runtime. PAA keeps four linked records: the declaration, evaluator evidence, the decision snapshot, and the autonomy event stream.
A task enters PAA when its boundary satisfies the instrumentation entry criterion.
Manual is the first on-spectrum position: the boundary is bounded and instrumented, position is governed, and the human performs the governed effect. Automation behind that boundary may be absent, disabled, or producing shadow outputs, none of that changes the position, because deployment is a separate axis.
The former partial-autonomy label that once sat between manual and hitl is retired. It no longer names a position, a display region, or an alias for manual, and nothing on the spectrum sits between those two.
Deployment, active, shadow, or disabled, is orthogonal to autonomy position. It is never an authority, a spectrum position, or a synonym for advisory.
One governed task has three independent lanes.
Autonomy position describes how much human oversight the task receives.
Evaluator maturity describes how reliably the work can be judged.
Worker capability describes the model, agent, workflow, or code performing the task.
Each lane moves only on its own evidence. Improving the worker does not earn autonomy, promoting the evaluator does not promote the task, and no lane inherits progress from another.
Behavioral and economic fitness are evidence dimensions used to assess a configuration across these lanes. They do not add an autonomy position or merge worker capability with evaluator maturity. A configuration can meet the task's behavioral bar and still be too expensive to operate at the proposed oversight.
The task declaration is static; the other three are produced at runtime against it. Full field contracts for all four are at /reference.
The implementation records economic measurements alongside evaluator verdicts. The
current evidence contract permits task-specific cost and configuration metadata in
payload, typed by the consumer's payload_schema. It does not
define top-level cost or worker fields, and runtime eligibility logic does not read cost.
Operators or consumer policies may use that evidence to keep the current configuration, choose a qualified replacement at unchanged authority, or inform a separate authority decision. The runtime does not execute model sweeps or select workers.
The versioned contract for the boundary, evaluators, deployment, position policy, promotion, and demotion.
One attributable evaluator verdict about one case or run.
The frozen evidence window behind one promotion or demotion decision, including exclusions and reasons.
The append-only motion history used to reconstruct approvals, rejections, and the current position.
Six areas connect the declaration to the runtime records. Exact fields live in the normative schema.
Define acceptable outcomes and relevant economic constraints alongside this declaration. A maximum cost per accepted result, a relative improvement target, or measurement without a hard gate are all possible policies. These are consumer operating requirements, not additional declaration fields. Read the fitness definitions.
Names the task and version, then defines the input, output, and action under control.
Names what gets judged, how the verdict is produced, which version produced it, and whether it can block the action.
Deployment says whether the candidate runs. Position says how much oversight the task requires.
Sets when evaluation runs at each autonomy position: before the action, after it, or offline.
Declares the evidence window, the proposed move, and whether an operator or automatic rule resolves it.
Declares when tighter oversight returns. An operator can demote immediately after a confirmed failure.
Validated example
This YAML is the validated refund_approval conformance fixture.
task: refund_approval
version: 1
description: Evaluate a refund request and decide to approve or escalate for human review.
boundary:
input: refund_request
output: decision
initial_position: hitl
deployment: active
evaluators:
- property: refund_policy_invariants
target: output
technique: deterministic
evaluation_basis:
kind: invariant
ref: refund_policy_invariants
epistemic_status: ground_truth
version: "1"
authority: blocking
- property: should_escalate
target: output
technique: classifier
evaluation_basis:
kind: reference_label
ref: should_escalate_reference_label
epistemic_status: ground_truth
version: "1"
authority: blocking
position_policy:
manual: offline
hitl: blocking
hotl: async
autonomous: offline
promotion:
from: hitl
to: hotl
report: refund_approval_promotion_report
window:
kind: cases
size: 200
execution: operator_approval
demotion:
from: hotl
to: hitl
trigger: chargebacks_or_complaints_exceed_bound
window:
kind: cases
size: 1