# Assessment Methodology

Framework version: 2.0.0-draft.1. Status: research proposal.

## Purpose and claim discipline

This methodology turns a security hypothesis into a bounded, reproducible assessment. It specifies artifacts and decision rules for investigating whether a defined contract can fail, how the failure occurs, and what conclusions the available evidence supports. It does not supply a universal certification threshold or establish that the proposed framework has already been empirically validated.

An assessment must separate the existence of a weakness, the frequency of an observed outcome under test conditions, the extent of demonstrated impact, and the generality of the proposed explanation. A single well-documented unauthorized state transition can warrant containment. It does not establish prevalence across deployments or a new attack family. Conversely, failure to reproduce an event is evidence about the tested conditions, not proof that the original report was false.

Required artifacts are a scope and authority record, system and threat model, property contracts, test protocol, execution record, analysis, falsification record, classification, and evidence package. Their detail should match the claim and consequences. An implementation-specific vulnerability need not demonstrate cross-model generality; a broad mechanistic claim requires evidence beyond one tailored example.

## Scope and authorization

Record the target configuration, owner, evaluator, purpose, permitted interfaces, test period, data handling constraints, and stop conditions before testing. Identify whether actions reach production systems, a staging deployment, a simulator, or stubs. Use synthetic records and isolated accounts wherever they can exercise the relevant contract. Permission to inspect a component does not imply permission to affect its users or connected services.

Establish containment before executing a candidate intervention. Describe network restrictions, tool substitutions, spending or resource limits, state snapshots, and recovery responsibilities as appropriate. A stop condition should specify who can stop the test, which observable event triggers it, and how residual work is cancelled. Tests involving physical actuation or human exposure require the additional controls described below; harmful live experiments are not part of this method.

An assessment of a nonadversarial hazard follows the same contract and evidence discipline but records the disturbance rather than assigning an attacker. The report distinguishes security findings, hazards, incidents, quality defects, and unresolved observations.

## Question and operational contract

Write the principal hypothesis before the confirmatory test. It should identify the actor or disturbance, available influence, target contract, expected failure event, and scope of inference. State a plausible alternative explanation and the observations that would weaken or refute the hypothesis.

Use P01–P23 as references, then operationalize the selected property. For example, P13 Action Integrity might require that the executor never commits a write outside the approved resource set for a named task. Specify what counts as approval, when it is checked, how scope changes, and which log or state readback demonstrates commitment. A model response saying it wrote a resource is not the same observation.

Define the outcome oracle before evaluating results. Prefer independently observable state, permission checks, or invariant violations over interpretation of prose. For semantic judgments, provide a coding rubric, examples of boundary cases, and an inconclusive category. If a model is used as a judge, freeze its configuration, validate it against independent annotations appropriate to the claim, and disclose disagreement and possible shared failure modes. A judge's label is evidence from that measurement process, not ground truth by definition.

## Design and controls

Establish a baseline under normal authorized operation. Use a matched negative control that preserves task content and relevant difficulty while removing the hypothesized malicious feature. Distinguish an authorized-function control, which verifies that legitimate operation remains possible, from a seeded-violation control, which verifies that the oracle detects a known contract failure in an isolated fixture. State which response each control should produce. All controls must remain contained and must not create real harm.

Document the intervention, what remains fixed, and unavoidable confounders. “Only one variable changed” is a design objective, not a statement to make when a payload also changes length, retrieval ranking, or task content. Use additional controls or bounded conclusions when these factors cannot be separated. Randomize or counterbalance order when order effects are plausible.

Exploratory search and confirmatory evaluation must be distinguished. Record the number and nature of payloads, prompts, or configurations explored, including unsuccessful attempts. Freeze a selected candidate before evaluating it on heldout tasks or fresh states. Continuing to adapt a payload against the same evaluation set changes the claim to performance under that adaptive search procedure.

## Units and sample planning

Define the independent experimental unit. A turn is not automatically independent of earlier turns; several outputs from one persistent session may form one unit. Multiple agents sharing a service or model state may be clustered. For human outcomes the unit may be a participant, team, or organization depending on assignment and dependence. State the unit of randomization, observation, and analysis when they differ.

Plan sample size around the claim, expected variability, smallest relevant effect or desired precision, dependence structure, and available resources. If information is insufficient for a justified confirmatory calculation, label the study exploratory and use a pilot to estimate design parameters. Do not adopt a fixed trial count merely because an earlier example used it. State stopping rules and how sequential monitoring or multiple comparisons will be handled.

Report denominators, exclusions, incomplete runs, timeouts, and missing observations. Distinguish invalid harness executions from genuine failures of system availability. Excluding inconvenient outcomes after seeing their condition can bias the result. Predefine invalidation rules and report sensitivity to disputed exclusions.

## Configuration and provenance

Record model identifiers and available revision details, sampling parameters, seed support, prompts or policy artifacts, context construction, tool schemas, permissions, retrieval snapshots, memory state, software versions, and relevant timing. Record unknown or provider-controlled variables explicitly. A public model name may not identify immutable behavior; include test dates and any available deployment revision.

Store exact inputs and outputs where permitted, along with hashes, timestamps, run identifiers, and execution logs. A seed can aid reproducibility without guaranteeing deterministic execution across platforms or provider changes. If state cannot be restored exactly, describe the reset approximation and test its adequacy. Record the effect of caches, concurrency, rate limits, and background updates where they can influence the outcome.

Protect secrets and personal information in the evidence package. Replace sensitive values with synthetic equivalents before testing when feasible. Redaction should preserve the evidence needed for the claim; if it does not, identify what an authorized reviewer must inspect privately.

## Execution and causal checks

Run baseline and intervention conditions under the planned protocol. Capture the earliest observable divergence, downstream transitions, and final outcome. Read back relevant state rather than relying solely on a tool's acknowledgement. Record refusals, partial actions, recoveries, and delayed effects as separate outcomes when the contract requires that distinction.

Use ablations to investigate necessary conditions: remove the asserted authority cue, disable the relevant persistence path, replace the suspect source, or interrupt a hypothesized feedback relation. These are tests of the explanation, not automatic fixes. An intervention that removes all system functionality cannot by itself show that a targeted control is effective.

Conventional static analysis or a direct access-control test may establish some implementation contracts without stochastic trials. Record the path and preconditions and verify relevant state transitions safely. Do not force every assessment into a language-model experiment merely because the surrounding product uses AI.

## Temporal and adaptive evaluation

For multi-turn cases retain the complete sequence and identify the independent session unit. Test whether earlier turns are necessary, whether order matters, and whether the effect persists after a documented reset. For L6 Memory & State Continuity, separate the write event, retained representation, later retrieval, and subsequent decision. An immediate response does not establish cross-session persistence.

For adaptive systems, record the initial state, update rule or service behavior available to the evaluator, authorized learning inputs, update schedule, and rollback mechanism. Compare matched update histories or replayable streams where possible. A snapshot test does not cover later adaptation. Holdout conditions must remain outside the adaptation process if they are intended to measure generalization.

For actions, identify the last effective intervention point and the point of external commitment. Test cancellation and revocation in a contained environment, including relevant races and retries. Report whether attempted, queued, executed, compensated, and irreversible outcomes are distinguishable in the evidence.

## Human and physical evaluation

Claims that presentation changed human decisions require human evidence. A narrower P19 Human Decision Integrity presentation or informed-authorization contract can be assessed through an inspectable mismatch, without claiming that anyone was deceived or changed a decision. Interface inspection can establish a mismatch or missing control without establishing its population effect. Label that narrower result accurately.

Human-subject research requires appropriate ethics review, informed consent, justified recruitment and sample planning, privacy protections, and debriefing where approved deception is involved. Use benign scenarios without actual financial, health, employment, or security consequences. Specify validated measurement instruments where appropriate and cite their sources in the individual study. Blind outcome coding when practical and account for repeated observations from the same participant. No universal human influence instrument is supplied here.

Physical-system tests should begin with simulation, recorded observations, hardware isolation, or nonhazardous fixtures. Document the gap between these conditions and deployment. Relevant practitioners must review containment and acceptance rules before any real-world actuation study. Do not expose people, animals, or operational infrastructure to a harmful condition to demonstrate reachability. A simulated contact or prohibited action remains a simulated outcome.

## Collective evaluation

For L9 Collective & Systemic Interaction, identify the collective contract and graph or subgraph to which it applies. Record participants, shared resources, scheduling, message semantics, initial states, and coupling. Measure the proposed propagation or amplification relative to a matched baseline, with its denominator and time window.

Vary relevant composition assumptions: participant count, topology, delays, shared dependencies, and local policies. A result that depends on one arrangement can still establish a vulnerability of that arrangement; it does not establish a universal property of multi-agent systems. Independent local behavior does not rule out a collective failure, while correlated failures from one shared service do not demonstrate agent-to-agent propagation.

Use relation ablation, shared-resource substitution, or alternative schedules to test the causal explanation. Permit group-level or subgraph loci when no single edge is responsible. Simulation assumptions and validation limits must accompany all systemic claims; modeled evidence is not equivalent to observation of a deployed population.

## Analysis and uncertainty

For each condition report valid units, observed contract violations, other outcomes, and uncertainty appropriate to the design. An observed proportion describes the tested distribution and procedure. It is not deployment likelihood unless the sampling and exposure assumptions justify that inference. Report differences or other prespecified effect estimates with uncertainty rather than only a significance label.

Account for clustering, adaptive selection, multiple comparisons, and incomplete observations where relevant. Use paired analysis when the design is paired. For small or sparse data, make the limits visible rather than presenting precise-looking percentages alone. If statistical modeling is used, record assumptions, diagnostics, and sensitivity analyses in the evidence package.

No observed failures in a finite test does not prove zero risk. A statistical bound, when reported, depends on independence, sampling, and model assumptions. Irreversible outcomes remain amenable to probabilistic analysis, but acceptable average performance does not compensate for a prohibited catastrophic event. Reachability, containment, intervention time, and frequency estimates can all matter; the applicable profile defines the decision rule.

## Falsification and classification

Actively test alternative explanations. Check whether the baseline exhibits the same violation, whether the intervention changed authorization, whether the oracle misread a harmless output, and whether hidden state or harness behavior explains the result. A mechanism-specific hypothesis can fail even while a real vulnerability remains. Document both outcomes.

Map entry points, failed contracts, propagation, and impact separately. Use the canonical L1–L9 names and stable M01–M12 and D1–D7 identifiers. Traversal alone does not establish a failed domain. Permit multiple causal failures, composite classification, ambiguity, and unrepresented cases. A forced primary layer can hide the most informative part of a finding.

Compare a proposed new family with the closest existing mechanisms and categories. New wording, a new target product, or a larger consequence does not by itself establish mechanistic novelty. Conversely, a narrow but real implementation weakness should not be rejected because it is not novel. Registry review of a class proposal is a different decision from remediation of a concrete finding.

## Evidence descriptors

The source framework's evidence labels are retained as descriptive facets rather than a compulsory cumulative ladder. Cross-model testing and independent replication answer different questions. A carefully controlled implementation finding can be strong without cross-model applicability.

| Descriptor | Evidence question |
| --- | --- |
| E0 Hypothesis | Is there an explicit testable claim and causal explanation? |
| E1 Single Observation | Is at least one relevant event documented? |
| E2 Reproduced Observation | Does the event recur under specified repeat conditions? |
| E3 Controlled Experiment | Do appropriate controls support attribution to the intervention? |
| E4 Cross-Condition Evidence | Does the stated claim hold under relevant condition changes? |
| E5 Cross-Model Evidence | Has the claim been examined across relevant model implementations? |
| E6 Independent Replication | Has a separate evaluator reproduced the claim with disclosed dependencies? |

For each facet record supported, unsupported, not_tested, inconclusive, or not_applicable, with evidence references and rationale. Unsupported means the cited evaluation does not support that facet, not that no vulnerability exists. A facet marked supported must identify the exact claim and tested scope. Independence includes authorship, analysis, data, and execution dependencies; reviewers should disclose which are shared.

Use assessment outcomes supported, refuted, inconclusive, quality_issue, hazard_only, or out_of_scope, with a reason and claim reference. These outcomes are separate from registry review status. Mixed findings can contain several claims with different outcomes. E0–E6 are not numerical confidence scores, and their identifiers must not be averaged.

## Severity and operating decisions

Describe severity through the consequences of a successful violation: affected assets or people, scope, privilege, duration, recoverability, and demonstrated versus potential harm. Describe likelihood through exposure, attacker opportunity, prerequisites, and available evidence. Describe confidence through the strength and limits of the causal and measurement evidence. Record reproducibility and tested coverage separately.

Reversibility is a contextual property of an action or impact. Record whether restoration, containment, or compensation is possible, by whom, at what cost, and within what time. Compensation does not necessarily undo disclosure or injury. A mixed sequence can include reversible internal changes and irreversible external effects.

An operational decision should cite a profile or named decision authority and explain why evidence warrants containment, remediation, further study, or bounded acceptance. Urgent containment can precede complete validation. Absence of a reviewed profile means the assessor must report unresolved acceptance criteria; it does not authorize inventing a universal score or claiming that the deployment passes OSAFIS.

## Evidence package and handoff

The package includes versioned scope, graph, contracts, threat capabilities, protocol, oracle, controls, sample rationale, configurations, raw or protected artifacts, execution accounting, analysis, falsification, classification, impact, and limitations. Give artifacts stable identifiers and integrity hashes where useful. State which material is public, restricted, unavailable, or destroyed under a retention rule.

The following design example is illustrative and unexecuted: a retrieval intervention is compared with matched benign material using a synthetic retained preference and a stubbed write tool. The protocol defines a session as the unit, verifies state reset, measures unauthorized preference commitment through readback, and separately records whether any proposed tool action would exceed scope. Trial counts and results are deliberately unspecified until sample planning and execution occur. No registry recognition follows from this example.

Submit concrete findings and class proposals to the registry as distinct entry kinds. The report should be useful to another evaluator without requiring private interpretation by its author. Responsible disclosure and evidence access follow Vulnerability Registry; profile-specific decision rules follow Domain Profiles.

## Validating the framework itself

Testing systems does not validate the classification framework. Evaluate that framework with a separate study: freeze definitions and coding guidance, construct a documented case sampling strategy, train coders on development cases, and reserve heldout cases for evaluation. Include conventional, semantic, temporal, human, physical, adaptive, and collective cases as the intended scope requires, including negative examples.

Collect independent classifications before adjudication. Record agreement and uncertainty for contract identification, domain mapping, mechanism mapping, and scope separately. Report ambiguity, unrepresented cases, and composite cases rather than treating them as coder errors by default. Analyze disagreements and assess whether proposed splits or merges improve useful distinctions on new cases. Measure assessor effort and whether the result supports actionable controls, not only label agreement.

Publish sampling limitations and avoid extrapolating from a convenient corpus to future completeness. Nine domains remain a working hypothesis. Revision is warranted when repeated boundary failures, missing contracts, or redundant categories undermine assessment utility. Changes require versioned migration rather than silent relabeling of earlier findings.
