Build PAA / Framework

Framework

A task declaration is not the runtime. PAA keeps four linked records: the declaration, evaluator evidence, the decision snapshot, and the autonomy event stream.

Recorded decision: spectrum entry

A task enters PAA when its boundary satisfies the instrumentation entry criterion.

Manual is the first on-spectrum position: the boundary is bounded and instrumented, position is governed, and the human performs the governed effect. Automation behind that boundary may be absent, disabled, or producing shadow outputs, none of that changes the position, because deployment is a separate axis.

The former partial-autonomy label that once sat between manual and hitl is retired. It no longer names a position, a display region, or an alias for manual, and nothing on the spectrum sits between those two.

Deployment, active, shadow, or disabled, is orthogonal to autonomy position. It is never an authority, a spectrum position, or a synonym for advisory.

Recorded decision: three independent lanes

One governed task has three independent lanes.

Autonomy position describes how much human oversight the task receives.

Evaluator maturity describes how reliably the work can be judged.

Worker capability describes the model, agent, workflow, or code performing the task.

Each lane moves only on its own evidence. Improving the worker does not earn autonomy, promoting the evaluator does not promote the task, and no lane inherits progress from another.

Three independent dimensions of a governed task Governed task: Task Governed task Autonomy position: Oversight Autonomy position Evaluator maturity: Evaluator Evaluator maturity Worker capability: Worker Worker capability
  • Autonomy position
  • Evaluator maturity
  • Worker capability
Authority, judgment, and capability can each change without implying movement in either of the other two.

Behavioral and economic fitness are evidence dimensions used to assess a configuration across these lanes. They do not add an autonomy position or merge worker capability with evaluator maturity. A configuration can meet the task's behavioral bar and still be too expensive to operate at the proposed oversight.

Four linked artifact layers

The task declaration is static; the other three are produced at runtime against it. Full field contracts for all four are at /reference.

The implementation records economic measurements alongside evaluator verdicts. The current evidence contract permits task-specific cost and configuration metadata in payload, typed by the consumer's payload_schema. It does not define top-level cost or worker fields, and runtime eligibility logic does not read cost.

Operators or consumer policies may use that evidence to keep the current configuration, choose a qualified replacement at unchanged authority, or inform a separate authority decision. The runtime does not execute model sweeps or select workers.

Task declaration

The versioned contract for the boundary, evaluators, deployment, position policy, promotion, and demotion.

Evidence record

One attributable evaluator verdict about one case or run.

Decision artifact

The frozen evidence window behind one promotion or demotion decision, including exclusions and reasons.

Autonomy event

The append-only motion history used to reconstruct approvals, rejections, and the current position.

Declaration architecture

Six areas connect the declaration to the runtime records. Exact fields live in the normative schema.

Define acceptable outcomes and relevant economic constraints alongside this declaration. A maximum cost per accepted result, a relative improvement target, or measurement without a hard gate are all possible policies. These are consumer operating requirements, not additional declaration fields. Read the fitness definitions.

Declaration identity and boundary

Names the task and version, then defines the input, output, and action under control.

Evaluator set

Names what gets judged, how the verdict is produced, which version produced it, and whether it can block the action.

Deployment and initial position

Deployment says whether the candidate runs. Position says how much oversight the task requires.

Top-level gate policy (placement)

Sets when evaluation runs at each autonomy position: before the action, after it, or offline.

Evidence-backed promotion

Declares the evidence window, the proposed move, and whether an operator or automatic rule resolves it.

Demotion

Declares when tighter oversight returns. An operator can demote immediately after a confirmed failure.

Validated example

refund_approval: the declaration in concrete form

This YAML is the validated refund_approval conformance fixture.

refund_approval.v1.yaml
task: refund_approval
version: 1
description: Evaluate a refund request and decide to approve or escalate for human review.

boundary:
  input: refund_request
  output: decision

initial_position: hitl
deployment: active

evaluators:
  - property: refund_policy_invariants
    target: output
    technique: deterministic
    evaluation_basis:
      kind: invariant
      ref: refund_policy_invariants
    epistemic_status: ground_truth
    version: "1"
    authority: blocking

  - property: should_escalate
    target: output
    technique: classifier
    evaluation_basis:
      kind: reference_label
      ref: should_escalate_reference_label
    epistemic_status: ground_truth
    version: "1"
    authority: blocking

position_policy:
  manual: offline
  hitl: blocking
  hotl: async
  autonomous: offline

promotion:
  from: hitl
  to: hotl
  report: refund_approval_promotion_report
  window:
    kind: cases
    size: 200
  execution: operator_approval

demotion:
  from: hotl
  to: hitl
  trigger: chargebacks_or_complaints_exceed_bound
  window:
    kind: cases
    size: 1

Canonical flows and use cases