OSAFIS Research Vision
OSAFIS proposes a common representation for analysing security failures in intelligent systems. Its purpose is to connect the security contract that failed to the component or relationship involved, the mechanism of influence, the evidence supporting the finding, and the consequences within a declared system boundary. Researchers can use this representation to compare findings. Engineers can use it to identify controls and assessment obligations. Neither use establishes that the representation is complete or that a system is secure.
The current specification is 2.0.0-draft.1. It defines a research method and a candidate vocabulary. Independent evaluation of classification reliability, assessment usefulness and coverage remains necessary. OSAFIS does not claim recognition as a standard, authority to certify systems, an operational public registry, or established industry adoption.
The problem being addressed
An intelligent system includes more than its learned model. A deployment may combine application code, training and retrieval data, observations, retained state, tool access, human approval and dependencies shared with other systems. Its security depends on the contracts between these functions as well as the integrity of each component. A valid credential can permit an operation that a user never requested. A retrieved document can provide useful evidence while having no authority to change the task. An accurate observation can contain text that the system must not treat as an instruction.
These distinctions require an account of authority and intended use. Conventional access control, software security and operational risk management remain part of that account. Existing AI security work already provides terminology, threat knowledge and mitigation guidance. The research opportunity for OSAFIS is to test whether a consistent contract-centred representation makes the relationship between these resources and a particular system easier to analyse.
The proposal focuses on identifiable failures and defensible evidence. A plausible attack description does not establish that a system permits it. A surprising output does not, by itself, establish a security vulnerability. A demonstrated failure in one configuration does not establish its prevalence across products. Each statement should retain the assumptions and observations that make it meaningful.
Scope and intended users
The intended scope includes learned and symbolic models, language and multimodal applications, retrieval systems, adaptive services, tool-using agents, embodied systems and interacting deployments. Applicability follows the functions present in the assessed system. A digital agent may observe and estimate an environment without a physical sensor. A predictive service may have no persistent task memory. A system marketed as autonomous may still contain human approval boundaries that determine its effective authority.
Security researchers need stable concepts, examples and refutable classification rules. System builders need security contracts with observable acceptance criteria and accountable control owners. Assessors need an evidence package that distinguishes what they observed from what they inferred. Domain specialists need a way to adapt tests to actual hazards, users and operating constraints. Affected people need their interests and avenues for challenge to be represented when an assessment makes claims about them.
The framework can describe safety-relevant failures, including accidental ones, where they concern a declared security contract or a material coupling with an intelligent system. It does not replace a domain safety assessment, clinical study, legal analysis or general account of social harms. Purely human organisational problems with no material connection to the intelligent system fall outside this scope. Unclear cases remain explicitly unresolved until the assessment boundary is justified.
The proposed representation
Nine provisional analytical domains describe security responsibilities. They are L1 Models & Computation, L2 Software & Infrastructure, L3 Data & Knowledge, L4 Perception & World Representation, L5 Interpretation & Objectives, L6 Memory & State Continuity, L7 Planning & Action, L8 Human–System Interaction and L9 Collective & Systemic Interaction. Their numbers identify them. They do not specify execution order, increasing intelligence, severity or expanding consequence.
Twenty-three property identifiers describe obligations that may apply in more than one domain. Twelve mechanism identifiers provide overlapping descriptors for adversarial influence. Seven cross-cutting dimensions prompt examination of identity and authority, governance, provenance, privacy and safety, change management, observability, and recovery. These counts describe the present vocabulary. Their optimality has not been demonstrated.
A separate graph records actual components, humans and shared resources, with typed relationships. A component can implement several domains. A common model should appear as a shared component when several agents actually depend on the same artifact, rather than as fictitious independent copies. Observations, information exchange and state synchronisation do not automatically transfer authority. Collective contracts can concern a group or feedback cycle without reducing to one defective pairwise connection.
Findings distinguish entry points, failed contracts, participating domains, propagation and impact. A downstream domain is not automatically vulnerable because an attack reaches it. Wide impact alone is not an L9 finding. Assessors may identify multiple causal failures or leave a classification ambiguous, without forcing an artificial primary layer.
Research principles
Definitions should support decisions. Each domain must identify a protected function or relationship and explain how it differs from its neighbours. Each property must specify an obligation that can be instantiated as a testable contract. A property name such as integrity or trust is insufficient without a subject, authorised change rule, relevant assumptions and observable violation.
Evidence should remain traceable. Reports should preserve configuration, versions, input provenance, test boundaries, outcomes, contrary evidence and restrictions on reproduction. Openness does not require publishing private data or a hazardous payload. Restricted evidence should have a declared access process and clearly stated consequences for independent verification.
Claims should remain proportionate. Classification coverage, attack detection, mitigation effectiveness and residual risk are separate questions. A framework can describe a failure it cannot automatically detect. A control can block a tested scenario without eliminating the broader class. An empirical study can support a bounded conclusion without proving absence of future failures.
Revisions should be reviewable. Disagreement is evidence about unclear boundaries, competing causal explanations or differing threat assumptions. The process should preserve this information. Changes to domain scope require a migration record, not silent reinterpretation of old findings. External mappings should preserve their source edition and distinguish a conceptual relationship from an equivalence claim.
Deliverables and their responsibilities
Foundational Concepts defines the unit of analysis and the terminology shared by the corpus. Security Layers defines the nine domains and their boundary rules. Security Properties defines the obligations used to construct security contracts. Threat Model records who can influence what, under which knowledge and access conditions. Attack Taxonomy describes the mechanism vocabulary and its limits.
Assessment Methodology defines how to scope, test, interpret and report an assessment. Domain Profiles applies the common method to candidate system classes without inventing sector-wide thresholds. Vulnerability Registry specifies evidence records and a proposed review process. Relationship to Existing Frameworks records source-qualified mappings. Versioning and Identifiers controls citation and migration. Future defines research questions and triggers for revision. This Vision document states the purpose and the claims the project is prepared to make.
The worked examples accompany the specification as unexecuted teaching cases. Machine-readable definitions and record validation support consistent documentation. They are not an attack scanner or evidence that a vulnerability exists. A website and presentation communicate this same material and should never imply greater maturity than the underlying evidence supports.
Evidence required for the next release
The architecture should be evaluated on a declared case population spanning different functions, lifecycle stages and consequence scopes. A development set can help clarify definitions. A separate held-out set, classified under frozen definitions by independent reviewers, can test whether those clarifications generalise. Reports should include uncertainty, disagreement, ambiguous cases, unrepresented cases and the effort required to produce a useful classification.
Comparative studies should test concrete questions, such as whether reviewers identify missing action-authorisation checks more consistently or trace a shared-resource failure more accurately. A comparison should use comparable information and task conditions. It should not score competitors by whether their documents happen to use the OSAFIS layer names.
Assessment usefulness requires its own evaluation. Teams should examine whether the method produces reproducible contracts and controls that address the observed failure while preserving authorised utility. They should track unsuccessful mitigation attempts and newly introduced failure modes. Human-facing studies require appropriate participant protections. Physical or irreversible effects require a safe test environment and explicit operational authority.
Governance and publication
The project needs a named maintainer, a documented review process, conflict-of-interest handling and an appeal route before it represents a registry as independently governed. A proposed role is not an appointed reviewer. A proposed publication licence is not an adopted licence. Until the owner makes those decisions, the material should describe the intended open-reference direction without granting unspecified rights or claiming institutional endorsement.
Success would mean that independent users can apply, challenge and improve the representation, and that its assessments support better security decisions in specified settings. Establishing that result requires evidence beyond completing the documents. This specification makes the commitments and tests explicit so the project can pursue that evidence.
Word document · Editable source
Foundational Concepts
OSAFIS 2.0.0-draft.1 | Research proposal | 7 September 2026
Purpose and status
OSAFIS proposes a vocabulary for describing security failures in intelligent systems and the evidence needed to assess them. Its unit of analysis is a deployed or specified system in an operational setting. A model, an application, a human approval process, and a network of cooperating services can each contribute to the same failure. A useful account must identify their actual relationships rather than attribute every consequence to the model.
The definitions below are conventions of this proposal. They do not establish that its categories are exhaustive, mutually exclusive, or empirically superior to other approaches. Their value depends on whether independent assessors can apply them consistently, produce useful tests, and identify actionable controls. Requirements stated here govern an assessment claiming to apply this draft; they are not claims that any implementation has satisfied them.
Intelligent systems and the assessment boundary
An intelligent system is a computational system that interprets inputs, produces judgments or outputs, or selects actions using learned or explicitly specified inference procedures. It may include models, conventional software, knowledge resources, observations, retained state, planning, tools, humans, and physical devices. It need not possess every function, maintain long term memory, or act autonomously. These characteristics determine assessment scope rather than eligibility by product label.
The system boundary identifies the components and relationships included in a claim. Its environment includes relevant actors, dependencies, resources, and conditions outside that boundary. An external service may be outside implementation control while remaining inside the causal account. Assessors must state who controls each element, which versions and configurations are examined, and which environmental assumptions support the claim. A result obtained from an isolated model endpoint does not automatically characterize a deployed workflow using that model.
The preferred representation is a graph of actual components, humans, shared services, and resources. Nodes may implement several security domains. Edges are typed as observation, information, instruction, authority delegation, action, state synchronization, feedback, or resource dependency. An edge can carry more than one explicitly identified relation. Receiving information does not itself grant authority. Shared nodes must remain visible: drawing a separate copy of one shared memory service for every agent would conceal a common dependency. Some failures concern a group or subgraph rather than one component or pairwise edge.
Security contracts and protected properties
A security property names a kind of obligation, such as Confidentiality or Objective Integrity. A security contract instantiates that obligation for a particular object or relationship. It states the protected object, authorized principals, permitted changes or effects, relevant operating conditions, and an observable criterion for violation. The contract must precede the evaluation result; defining it after seeing an unwanted output invites circular classification.
For example, “the system must be safe” is insufficient. A more assessable contract could require that a document review service disclose records only to the requesting account, use retrieved documents as evidence rather than as authority to change the task, and obtain the designated approver's authorization before sharing a report. Each clause supplies a separate potential failure criterion. The policy owner must specify what counts as an account match, valid task change, and approval.
P01–P23 in Security Properties provide reusable property definitions. They are neither a mandatory checklist of universally applicable functions nor an assertion that all obligations have been identified. Several properties can apply to one contract. A more specific property normally provides a more informative label than generic Integrity; independent failures should still be recorded separately. A property identifier is a reference to a definition, not evidence of protection.
Authorized objectives and contested goals
An authorized objective is a task or outcome established by a principal entitled to set it, within the governing constraints and delegation applicable to that system. Its source must be recorded: for example, a deployment policy issued by the responsible organization, a valid request from an authenticated user, and the limited authority granted to a service acting on that request. The precedence among these sources is a deployment decision that must be made explicit. The framework does not assume a universal ordering of all developers, operators, users, affected people, and external authorities.
Authorization concerns both entitlement and scope. A genuine user's request does not necessarily authorize every means of accomplishing it or every effect on another person's resources. Retrieved content, remembered statements, and another agent's message cannot change an objective merely by claiming urgency or authority. Conversely, a legitimate authorized change is not an integrity failure simply because the task changes.
Goals can conflict or be underspecified. The assessment should record the competing requirements, the designated decision owner, and the permitted resolution procedure. Where a consequential action depends on unresolved authority, the contract should define a pause, escalation, or restricted fallback. Assessors should report an unresolved governance assumption when no such rule exists, rather than inventing the “true” objective. Objective Integrity protects the established authorization relation; it does not certify that an authorized objective is ethical, lawful, or desirable. Those questions require their own substantive criteria and accountable review.
Threats vulnerabilities attacks and incidents
A threat is an actor capability, condition, event, or circumstance that could defeat a security contract. A threat model identifies who can influence what, what they know, what access they possess, and what constraints limit them. It can also describe nonadversarial disturbances relevant to the same contract. An attacker is not presumed merely because a failure occurred.
A vulnerability is a weakness in a component, configuration, process, or relationship that permits a security contract to be violated under stated conditions. Evidence must connect the weakness to the violation or establish a sufficiently supported feasible path. Unusual text, an incorrect answer, and a dangerous possibility are not by themselves vulnerability demonstrations. A quality defect concerns failure against a performance expectation; it can also be a security weakness when it defeats a defined protection. A hazard is a condition with the potential to cause harm. An incident is an actual event with security relevance or impact. These categories can overlap without becoming synonyms.
An attack technique is a method of exercising influence to exploit a weakness. An exploit is a concrete execution or implementation of such a technique against a specified target. A mechanism describes how the influence operates at a useful level of abstraction. Attack Taxonomy separates mechanisms, families, techniques, and modifiers; the M01–M12 identifiers retain their canonical meanings. Naming a mechanism does not establish the existence, novelty, or exploitability of a vulnerability.
A benign fault can demonstrate a contract failure without establishing an adversarial exploit. For example, a delayed authorization update can expose a stale permission check even if nobody intentionally caused the delay. The assessment must distinguish the observed trigger from any hypothesized attacker capability to reproduce it. This permits security and safety evidence to inform one another without fabricating an adversary.
Domains dimensions lifecycle and impact
The legacy term “layer” is retained in identifiers L1–L9, but denotes a provisional analytical security domain. A domain groups protected objects, functions, or relationships with related contracts. The domains are Models & Computation; Software & Infrastructure; Data & Knowledge; Perception & World Representation; Interpretation & Objectives; Memory & State Continuity; Planning & Action; Human–System Interaction; and Collective & Systemic Interaction.
These domains are not execution steps, severity levels, or a progression of consequence scope. Nine is a working partition to be evaluated. Functional applicability matters: a text based agent can estimate a changing digital environment, while a multimodal application may have no retained operational memory. An assessor may assign several domains to a component or leave a classification unresolved with a reason. Security Layers gives boundaries and discriminating questions.
A cross cutting dimension is a perspective that can inform contracts in several domains. D1 is Identity, Trust & Authorization; D2 Governance & Accountability; D3 Supply Chain & Provenance; D4 Privacy & Safety; D5 Lifecycle & Change Management; D6 Observability, Logging & Auditability; and D7 Resilience & Recovery. Applicability and evidence vary by context. Dimensions neither replace domain classification nor constitute additional layers.
Lifecycle phase and impact scope are separate descriptors. A weakness introduced during training may be triggered during operation and discovered during retirement. Its consequences may affect an individual, an organization, interconnected services, or a wider population. Large impact does not, by itself, establish L9: that domain requires a distinct collective interaction contract and a specified coupling mechanism. A compromised shared model can have broad consequences while its demonstrated failed contract remains in L1.
Causal paths and composition
An attack path connects an initial influence to a security consequence through observed or hypothesized transitions. A causal failure path is the broader term when malicious intent is absent or unknown. The record should separate entry point, failed contracts, intermediate propagation, controls encountered, and resulting impact. An edge traversed without violating its contract is part of propagation, not an additional vulnerability.
Paths need not be linear. Feedback loops, delayed state, parallel requests, and shared dependencies may require a subgraph with a timeline. Several independent contract failures can contribute to one incident. Selecting one primary domain for indexing is optional and must not suppress other supported failures. Conversely, applying several labels to the same obligation should not multiply the number of findings.
Consider an illustrative, unexecuted scenario: a retrieved report contains a false instruction to change the recipient of a summary. The report enters as information; the interpreter treats it as authority; a proposed send operation passes an inadequate recipient check. The entry point is the report. The potential instruction authority failure is in L5, and the independently inadequate action check is in L7. Merely retrieving and conveying the report does not prove an L3 failure. A valid recipient check would interrupt the path even if the interpretation failure remained.
Evidence and assessment claims
Evidence consists of observations, artifacts, and reasoning supporting a specified claim. Sources may include inspection, controlled experiments, incident records, reproductions, and independent replication. Each supports different conclusions. An observed failure may establish that a path is possible without estimating its frequency; a plausible architectural argument may justify testing without demonstrating exploitation. Evidence must preserve the relevant configuration, inputs, state, expected condition, actual observation, and limitations sufficiently for independent review where feasible.
An oracle is a rule or observation used to decide whether a contract was violated. A test should state its oracle, baseline, threat capabilities, negative controls, and stopping conditions. Negative controls help distinguish the proposed mechanism from ordinary task variability or an unrelated weakness. Instrumentation should observe decisions and effects needed for the claim without assuming access to reliable private reasoning. Tests proposed in this corpus are illustrative and unexecuted unless accompanied by a separate execution record.
Confidence describes how strongly the evidence supports a claim. Reproducibility concerns recurrence under specified conditions; repeatability in one setup is narrower than independent replication. Coverage describes the portions of a defined threat and operating space examined. None is interchangeable with severity, and no finite evaluation establishes zero risk. Assessment Methodology specifies the study record; Vulnerability Registry specifies the claim and review record.
Risk severity reversibility and controls
Risk concerns the possibility and consequences of a security loss in a specified context and time horizon. It depends on exposure, threat capability, system behavior, controls, and uncertainty. Severity describes the consequence of a specified successful failure under stated assumptions. A severe possible consequence may have weak supporting evidence or limited exposure. These facts must remain separately visible; this draft does not prescribe a universal scalar or multiplication of ordinal labels.
Reversibility concerns an action's effects under particular recovery conditions. Restoration, containment, and compensation are distinct: restoring a file does not undo its disclosure, and paying compensation does not restore an injured person. Recovery can be partial and time limited. Assessors should identify the point at which the relevant effect becomes irreversible, who can intervene before it, and what remains recoverable afterward.
Probabilistic evidence remains meaningful for irreversible harm. Irreversibility changes acceptable decision rules and the need for containment; it does not make likelihood meaningless or imply that estimating a rate accepts the harm. Rare or catastrophic outcomes require particular caution about sparse observations, dependence, distribution change, and unobserved paths. Tests should use simulation, inert substitutes, or contained environments where actual effects would be unacceptable.
A security control is a technical, procedural, organizational, or operational measure intended to preserve a contract or detect and limit its violation. A mitigation reduces exposure, likelihood, propagation, or impact without necessarily eliminating the weakness. Controls need their own assumptions and verification; a documented approval requirement is not evidence that execution is gated by approval. Residual risk and displaced failure paths remain part of the assessment.
Evolution and citation
An emerging vulnerability class is a candidate grouping whose protected obligation or mechanism is not adequately represented by existing definitions. A proposal should explain the classification difficulty, compare neighboring categories, present evidence, and allow independent challenge. Ambiguous cases are useful evidence about the framework's boundaries; they should not be forced into a familiar label to preserve apparent completeness.
Every classification must cite the framework version and stable identifiers. Versioning and Identifiers governs changes and migration. The source corpus labeled 1.0 supplies the historical baseline; that label does not establish public release. This revision changes domain boundaries and therefore requires human review of affected prior classifications. Scientific value depends on preserving what the evidence actually supports, including uncertainty, rather than on preserving the appearance of a complete taxonomy.
Word document · Editable source
Security Layers
OSAFIS 2.0.0-draft.1 | Research proposal | 7 September 2026
Analytical purpose
The nine OSAFIS layer identifiers organize analysis by the object, function, or relationship whose security contract may fail. In this revision, “layer” means a provisional analytical domain. The identifiers do not describe a processing pipeline, an implementation stack, levels of intelligence, or increasing consequence scope. A model artifact failure can affect many organizations; a collective interaction failure can remain confined to a small test environment.
Nine is a working hypothesis about useful distinctions. A proposed partition should earn its place through assessor agreement, explanatory value, and the ability to guide discriminating tests. This document does not establish that nine domains are necessary, sufficient, optimal, or future complete. Overlap between components is expected; overlap between definitions requires careful identification of the particular contract under examination.
The domain model complements three separate questions: which property is protected, how influence operates, and what consequences follow. Property identifiers P01–P23 answer the first; mechanisms M01–M12 in Attack Taxonomy address the second. Consequence scope, lifecycle phase, uncertainty, and recoverability are recorded independently. Cross cutting dimensions D1–D7 remain perspectives on these analyses rather than additional domains.
Representing the system before classifying failures
An assessment begins with a graph of actual components, humans, shared services, and resources. Assign domains to the functions performed by each node and to relevant relationship contracts. One service can implement retrieval, persistent state, inference, and action dispatch; it therefore need not have one exclusive domain label. Conversely, one domain contract can span several components.
Relations are typed as observation, information, instruction, authority delegation, action, state synchronization, feedback, or resource dependency. Document the payload, applicable identities, trust assumptions, and permitted transformations where they affect the claim. A status message is ordinarily an information edge; an instruction can ask for work without conferring new authority. Delegation exists only where some defined authority is transferred or exercised on another principal's behalf.
Shared services and resources are represented as shared nodes. Copies of a nine layer stack per agent can obscure a common model, datastore, budget, or controller. Group and subgraph contracts are also admissible: a coordination invariant may concern all active participants and cannot always be reduced to independent pairwise checks. Record observed and hypothesized links distinctly.
For every finding, separate the attack entry point, failed contracts, propagation, and impact. A traversed domain has not necessarily failed. Multiple failures can be causally necessary, and classification can remain unresolved pending evidence. A primary domain may support indexing, but it is not a rule that the first influenced component owns every subsequent violation.
L1 Models and Computation
Canonical display name: Models & Computation.
L1 protects executable model parameters, inference rules, and the agreed model computation. Its contract identifies the approved model artifact or rule set, permissible transformations, relevant computation settings, and the conditions under which the implementation is expected to realize that computation. This applies to learned and explicitly programmed inference systems. Architecture, weights, executable adapters, and computation affecting model configuration can be relevant protected objects.
The boundary with L2 is the distinction between the computation being specified and the software or infrastructure executing and serving it. Unauthorized replacement of approved weights is an L1 artifact integrity failure. A serving process that allows an unauthorized caller to replace files also has an L2 access control failure if that contract is independently defeated. Training data belongs to L3 as an information resource; the resulting model's failure to meet a defined security behavior may warrant an L1 finding, with the data influence recorded as the causal entry path. Merely producing an incorrect answer does not establish that the model computation contract failed.
An illustrative, unexecuted case concerns a deployment intended to load one approved model and adapter combination. A mutable alias resolves to an unapproved adapter after startup, although the base model remains unchanged. The protected obligation is to execute the approved combination, not a general promise of correct outputs. An assessment can compare the approved manifest with the artifacts actually loaded and exercise the alias change in an isolated environment. A negative control uses a permitted version transition with the required authorization and verification.
The assessable question is whether execution remains bound to approved computational artifacts and changes. Candidate controls include verified loading, immutable version references, explicit transformation records, and checks at reload boundaries. Their effectiveness depends on where verification occurs and whether subsequent mutation remains possible. A successful initial verification alone does not establish lifetime integrity.
L2 Software and Infrastructure
Canonical display name: Software & Infrastructure.
L2 protects implementation, runtime, access control, isolation, and infrastructure contracts. Relevant objects include service endpoints, credentials, processes, deployment configuration, dependencies, networks, storage access, and execution environments. The contract specifies who may invoke or modify which functions, how tenants and privileges are separated, and the service or resource bounds required for authorized operation.
The presence of a model does not change the ownership of a conventional session mix up or exposed management endpoint. L2 covers the mechanism enforcing an application's access boundary. L7 covers the task specific authority and effects of actions performed through that mechanism. A tool may authenticate a service correctly while that service requests an operation outside the user's task; this distinguishes a valid infrastructure check from a failed action contract. L1 remains responsible for approved model computation, while L2 covers the serving implementation and isolation supporting it.
An illustrative, unexecuted failure occurs when two agent sessions share a process cache whose key omits tenant identity. One session retrieves another tenant's response. The failed contract is tenant separation, even if the response contains model generated text and the visible impact is confidentiality loss. A contained test can use synthetic accounts and distinct marker data, alternate requests, and inspect whether account boundaries survive cache reuse. Repeating the sequence with caching disabled can help localize the mechanism without proving that every cache path is safe.
Assessment asks whether unauthorized principals or execution contexts can reach protected resources under the stated deployment configuration. Candidate controls include complete tenant binding, scoped service identities, isolation, and resource limits. Logging supplies evidence only if it captures the relevant identity and boundary decision. Broad downstream harm does not move the original infrastructure failure into L9.
L3 Data and Knowledge
Canonical display name: Data & Knowledge.
L3 protects information resources and the conditions under which they are supplied as knowledge. Its contract covers authorized changes, access, provenance, source status, and any declared requirements for freshness or verification. Training corpora, retrieval indexes, documents, metadata, reference databases, and tool supplied information can fall within this domain. The obligation is not that every statement is true; it is that the resource satisfies the explicitly required handling and epistemic conditions.
L3 differs from L5 at the use of information as instruction or objective authority. A correctly labeled untrusted document can carry hostile language without violating L3. If the interpreter elevates that language to instruction authority, the failed contract is in L5. L4 concerns acquisition and estimation of an environment state; L3 concerns information offered or maintained as a resource. L6 concerns information functioning as retained operational state. A database can implement both a knowledge collection and user memory; classify the contract, not the storage technology.
In an illustrative, unexecuted example, a supplier bulletin index drops a source identifier during ingestion. A later result is presented as coming from an approved publisher although its content originated elsewhere. The proposed L3 failure is loss of the source binding required by the index contract. A test can ingest distinct synthetic sources, inspect their transformed records, and compare retrieved provenance with the ingestion record. A negative control preserves the binding through the same transformations.
The assessable question is whether information crosses collection, transformation, retrieval, and disclosure boundaries with its required permissions and provenance intact. Candidate controls include source binding, transformation history, separation of verification status from content, and access checks at retrieval. A source signature can support an origin claim without establishing the truth or suitability of the source's statements.
L4 Perception and World Representation
Canonical display name: Perception & World Representation.
L4 protects observation and the estimation of an operational environment. Its contract specifies which features are observed, how observations are associated with entities and times, how uncertainty and stale state are handled, and when reobservation is required before a dependent decision. It includes physical sensors and digital observation, such as an agent inspecting a browser page or a service state. Applicability follows function: a text representation can be an observation of a changing environment.
The distinction from L3 is functional. A reference document describes information; a current view used to locate a live interface element estimates the state in which an action will occur. The same bytes can serve either role. L6 protects retention and authorized state updates, whereas L4 protects whether the represented environment remains adequately grounded. An accurately stored but obsolete belief can satisfy storage integrity while violating a required freshness contract. L7 covers the subsequent action validation and effect.
An illustrative, unexecuted case uses a simulated administration page. After an agent observes a selected test resource, the page updates and the selected row changes. The agent continues to represent the earlier resource as selected. An assessment can compare timestamped observations and the agent's externally inspectable state representation with the simulator's known state. The violation oracle is a declared requirement to refresh selection identity before a consequential operation. A negative control holds the page stable or forces reobservation.
Candidate controls include observation timestamps, entity binding, explicit uncertainty, consistency checks, and mandatory refresh conditions. No assumption is made that physical input is wholly unauthenticable or that authenticated input is accurate. Origin, sensor condition, interpretation, and environmental truth are different claims. Simulation results must identify the gap between the simulated environment and deployment. The expanded digital scope is a 2.0 migration issue for historical L4 and P22 classifications.
L5 Interpretation and Objectives
Canonical display name: Interpretation & Objectives.
L5 protects the assignment of meaning, instruction authority, and authorized objectives. Its contract defines which principals may establish or revise the task, the precedence and scope of their instructions, how quoted or retrieved material is treated, and how ambiguity or conflict is resolved. Applicable sources may include deployment policy, a valid user request, and a bounded delegation. Their precedence is specified by the system's accountable policy owner rather than presumed by the framework.
L5 is not a domain for every unwanted output. If the objective remains “assess this transaction” while an evidential threshold is applied incorrectly, the immediate question is Decision Integrity at the component performing that judgment. If the chosen strategy violates a sequencing constraint while the objective is unchanged, L7 Planning Integrity is more informative. L3 owns source handling; L5 owns elevation of source content into instruction authority. L8 examines what a person is led to understand or authorize.
An illustrative, unexecuted case gives a review service an authorized task to summarize a report for a named recipient. A passage inside the report claims that a new policy requires sending it elsewhere. The L5 oracle compares the task adopted after reading the passage with the authorized objective record and the contract forbidding task changes from report content. A control passage containing the same recipient information as an ordinary factual statement helps distinguish a directive from a content effect. Execution of a send is not needed to establish a supported objective change.
Assessment asks whether a source without task setting authority can change the governing instruction or objective. Candidate controls include explicit source roles, constrained task state updates, and conflict escalation. If authorized stakeholders disagree and no resolution rule exists, the assessor records an unresolved authority contract. An observed consequential disagreement cannot be resolved merely by declaring one interpretation “aligned.”
L6 Memory and State Continuity
Canonical display name: Memory & State Continuity.
L6 protects retained operational state and continuity across turns, sessions, interruptions, or resumed tasks. Its contract identifies who may create, modify, associate, read, expire, or delete state; which task or principal it belongs to; and the conditions under which it may influence future work. Conversation summaries, user preferences, checkpoints, pending authorizations, and persistent task records can qualify. Persistence is functional and may be brief; an in memory checkpoint can be security relevant.
L3 protects information as a resource; L6 protects information as continuity of a particular operation or relationship. L5 determines whether remembered text can legitimately function as an instruction. L4 assesses the accuracy and freshness of a remembered world representation. Temporal Integrity spans domains and does not automatically imply L6: an expiring tool credential can fail its time condition without any memory subsystem being defective.
In an illustrative, unexecuted example, a task summary records “user approved sharing” while omitting that approval covered only a particular synthetic report. On resumption, the approval is associated with a different report. The state continuity contract requires preservation of the authorization's object and scope. A contained assessment compares the precheckpoint authority record, serialized summary, restored state, and downstream permission request. A negative control resumes a task with all required bindings preserved.
The assessable question is whether retained state preserves its authorization, identity, provenance, and validity conditions through transformation and reuse. Candidate controls include structured scope binding, versioned state, expiry, explicit invalidation, and revalidation on resume. Deleting one visible memory record does not demonstrate complete revocation if derived summaries or other stores remain able to reintroduce it. Assessment should identify the stores actually examined and the limits of that claim.
L7 Planning and Action
Canonical display name: Planning & Action.
L7 protects strategies, capability selection, delegation, action validation, and execution. Its contract binds an authorized objective to permissible means, resources, recipients, sequence constraints, budgets, and external effects. It applies to digital operations and physical actuation. A relevant action may be a message, a file modification, a service request, or movement of a device; describing an operation as a tool call does not settle its authority.
L5 protects what the system is authorized to accomplish. L7 protects how it attempts and executes that task. L2 may authenticate a caller and expose a capability without establishing that a particular use is allowed. L8 covers the human's understanding and usable authorization interface. L9 applies when a separate collective interaction obligation is involved, not simply because one agent delegates to another. Within L7, P12 concerns capability use selection, P13 realized effect, P14 intervention, P15 strategy, and P16 delegated scope.
An illustrative, unexecuted case authorizes updating two test records only if both updates can be committed together. An agent constructs a plan to commit the first and then independently attempt the second. The plan violates the required transaction constraint before an external effect occurs. The oracle compares the inspectable plan or scheduled operations with the declared all or nothing condition. A simulated backend can then determine whether execution enforcement blocks the partial commit. A compliant transactional plan is a negative control.
Assessment asks whether permissible operations remain permissible as a sequence and at execution time. Candidate controls include constrained delegation, effect level checks, budget accounting, revalidation after delay, and intervention paths with measured deadlines. A stop button that cannot prevent the relevant irreversible effect within its declared timing requirement does not satisfy that controllability contract. Actual hazardous actions must be replaced with contained substitutes during evaluation.
L8 Human System Interaction
Canonical display name: Human–System Interaction.
L8 protects human understanding, informed authorization, oversight, and intervention in a system relationship. Its contract identifies the human's role, material information needed for a decision, claims the interface is allowed to make, and the conditions for meaningful approval or control. It can cover visible uncertainty, provenance, scope of proposed actions, and the correspondence between a displayed request and the operation subsequently authorized.
L8 does not assume that a model has human emotions or intentions. Nor does it classify every persuasive or inconvenient output as a vulnerability. A security relevant obligation and material misleading influence must be identified. L5 concerns the system's interpretation of authority; L8 concerns the person's understanding of the request or result. L7 owns action enforcement. An interface can misrepresent the proposed scope even if a backend correctly implements the approval token it receives.
An illustrative, unexecuted case presents approval for “share this summary” while the bound operation includes its full confidential attachments. In a synthetic workflow, assessment can compare the displayed scope, generated approval record, and planned external payload. The mismatch is inspectable without exposing real data or involving unsuspecting people. Establishing how a presentation changes human decisions requires a separate study design; interface inspection alone does not demonstrate that users were deceived.
Candidate controls include specific effect previews, explicit recipients and data scope, truthful provenance, and accessible intervention. Human subject studies require appropriate consent, review, debriefing, and data protections. Assessment must distinguish a design defect, evidence of materially misleading communication, and observed human impact. The presence of a person in the workflow does not itself demonstrate effective oversight or transfer responsibility away from the system's other control owners.
L9 Collective and Systemic Interaction
Canonical display name: Collective & Systemic Interaction.
L9 protects a distinct contract governing interactions among participants or dependencies. A finding requires a demonstrated or hypothesized coupling mechanism, such as feedback, shared resource contention, state synchronization, coordinated allocation, or correlated dependence. The contract may concern a group limit, consistency condition, containment boundary, or recovery behavior. Broad public or societal impact is recorded separately and does not suffice for this domain.
An L2 outage in a shared service can affect every dependent agent without demonstrating a separate L9 vulnerability. L9 becomes relevant when their interaction defeats an independently specified collective obligation. Similarly, a delegation scope violation is ordinarily L7; multiple participants do not automatically create a collective failure. A valid L9 analysis can identify a group failure even when each participant complies with its local contract, provided the composition contract and coupling are explicit.
Consider an illustrative, unexecuted simulation in which several schedulers share a limited resource. Each scheduler respects its local retry limit, but delayed observations cause their retries to synchronize, preventing the group from recovering within its contracted interval. The hypothesis concerns a feedback and shared dependency subgraph. Assessment compares recovery time and total concurrent requests with the group contract, then varies synchronization or introduces independent scheduling as a negative control. Merely counting many requests would not establish the proposed mechanism.
Candidate controls include shared admission rules, coordinated budgets, feedback damping, and isolation of failure domains. They must be evaluated at the relevant group scale and under the assumed communication delays. A small simulation supports claims about that model and tested conditions; it does not establish behavior at societal scale. If no collective contract or coupling can be specified, record the broad impact and leave L9 unassigned.
Classification and migration practice
An assessor should write the failed obligation in plain language before choosing a domain. Next, locate its actual component or relationship, identify the property and mechanism, and distinguish evidence of violation from mere propagation. Compare the closest neighboring domain using the boundaries above. Record independent failed contracts separately while retaining their common causal path. When the available observations cannot distinguish two explanations, preserve both hypotheses and specify the next discriminating test.
The source 1.0 labels map to 2.0.0-draft.1 identifiers as follows. This mapping preserves references; it does not imply unchanged scope.
The source's consequence ordering, automatic authority on every interagent edge, and exclusive emphasis on one primary layer are not retained. Historical findings require human reclassification where those assumptions affected the result. Versioning and Identifiers records the migration policy; Threat Model supplies scoped assumptions; Security Properties supplies obligations; and Assessment Methodology supplies test and evidence requirements. These interfaces make the domain model usable while leaving its scientific adequacy open to evaluation.
Word document · Editable source
Security Properties
OSAFIS 2.0.0-draft.1 | Research proposal | 7 September 2026
What a property establishes
A security property names an obligation that an assessment can instantiate for a protected object, function, or relationship. An attack mechanism describes how influence operates; a domain locates the relevant contract; a property states what that contract protects. These axes are related but not interchangeable. A malicious document may reach a system without violating any property, and the same property can fail through several mechanisms in several domains.
The twenty three identifiers below retain the assignments and labels in the source Versioning and Identifiers. They are proposed analytical conventions, not evidence that a deployed system satisfies them. The source Security Properties used numbered section headings that do not match all permanent identifiers. Citations must use P identifiers and a framework version rather than infer identity from section position. In particular, P14 is Controllability, P15 Planning Integrity, P17 Temporal Integrity, and P18 Decision Integrity.
Instantiating and testing a property
For each applicable property, specify the protected object or relationship, authorized principals, permitted operations, governing conditions, time horizon, and observable violation criterion. Identify the policy owner and evidence that the policy applies. “The output was harmful” does not specify the obligation, and “the system remained secure” is not an oracle. Contracts may include categorical prohibitions, limits, thresholds, or conditional obligations; their values require justification in the assessed setting.
An oracle decides whether the specified condition holds. It may compare authorization records with effects, inspect source bindings, evaluate a declared rule against controlled facts, or use a reviewed rubric for meaning and presentation. The oracle's limitations must be reported, including disagreements between human adjudicators. Model generated explanations can be observations but should not be assumed to reveal the actual causal process or supply independent ground truth.
The examples and tests below are illustrative and unexecuted. They provide possible assessment designs, not experimental results, validated vulnerabilities, or universal controls. Tests should use synthetic data and contained effects where practicable. Each execution record must state baseline behavior, manipulated condition, negative controls, observed outcome, configuration, and uncertainty. A property can be inapplicable, untested, satisfied within examined conditions, violated under stated conditions, or unresolved. These states must not be collapsed into a binary assertion of system security.
P01 Confidentiality
Confidentiality requires that protected information be disclosed only to principals and channels allowed by the applicable policy. The contract must identify the information, recipients, transformations, and outputs in scope. Instructions, model artifacts, and internal traces are confidential only when a policy actually protects them; internal placement alone is not a secrecy rule. Disclosure can occur through text, retrieval, memory, tool payloads, or other observable channels.
The violation oracle establishes that an unauthorized recipient obtained protected information or a prohibited inference under the declared leakage criterion. A contained test can place a synthetic secret in one account and examine another account's outputs, with an authorized recipient as a negative control. A guessed marker is weak evidence unless its origin can be distinguished from chance or permitted knowledge. P01 concerns disclosure; P02 concerns unauthorized alteration, and P03 concerns service access. One incident can independently involve all three.
P02 Integrity
Integrity requires that protected information, configuration, computation, or state be created and changed only through permitted transformations and by authorized actors. “Trustworthy” must be translated into a specified invariant, such as an approved artifact digest, valid update authorization, or preserved account binding. The generic property remains useful for objects whose obligation is not better expressed by a specialized integrity property.
The oracle compares the protected object's observed state or transition with the authorized version and change rules. A model artifact replacement test can inspect the actual loaded artifact rather than infer replacement solely from changed answers. A permitted update exercises the negative case. Use P07 for the particular semantics of retained operational memory, P04 for instruction authority, and other specific properties where they provide the precise obligation. Adding P02 to every specialized integrity finding does not establish an additional failure or increase its severity.
P03 Availability
Availability requires delivery of specified functionality to authorized users within declared service, resource, or timing bounds. Those bounds should identify the workload and environmental assumptions, including admission controls and dependencies. A refusal required by policy is not a denial of authorized service. An unbounded promise to answer every request is not a workable contract.
The oracle measures the specified service outcome, such as completion, latency, or admitted workload, under a controlled condition. A contained test might introduce bounded recursive work and measure whether unrelated authorized requests remain within the service contract. A workload of equivalent permitted cost helps distinguish the proposed amplification mechanism from ordinary capacity limits. P03 concerns the ability to obtain service; P14 concerns the ability of an authorized party to constrain or stop operation. A continuously running service can satisfy one while violating the other.
P04 Instruction Integrity
Instruction Integrity requires that authorized instructions preserve their meaning, authority, scope, and precedence through the processing path. The deployment must identify who can issue each instruction class and how conflicts are resolved. Retrieved passages, tool responses, and quoted speech are not promoted to a privileged instruction role merely because their wording resembles an order.
The oracle compares the accepted instruction set and its observable effects with the governing source and precedence rules. An illustrative test supplies a document containing a purported policy override and inspects whether a protected task constraint is displaced. A valid instruction from the authorized source provides a negative control against indiscriminate refusal. P11 concerns meaning preservation more generally; P04 specifically concerns instructions and their authority. P09 concerns the resulting authorized objective. Record both only when evidence establishes the instruction violation and the objective change, rather than assuming one from the other.
P05 Context Integrity
Context Integrity requires that the material assembled to frame a task preserve the required membership, attribution, ordering, and qualification of relevant information. The contract can specify which history, assumptions, observations, or evidence must accompany a decision and what contextual omissions are prohibited. It does not require every possible fact to fit in a model's context.
The oracle compares the assembled decision context with that contract. An illustrative test introduces a summarization step that retains an approval statement while dropping its limiting condition, then checks whether the limitation remains available in the decision context. A summary retaining both clauses is a negative control. P06 concerns the underlying knowledge resource; P07 concerns retained state across operations. P05 concerns the task frame actually supplied for the current decision. If a required qualifier was already lost in persistent memory, the memory failure and its contextual propagation should be distinguished.
P06 Knowledge Integrity
Knowledge Integrity requires that information used as knowledge retain its required provenance, verification status, and conditions of use. It does not assert universal truth. A system may legitimately use uncertain or conflicting sources if the contract preserves those qualifications and does not present them as verified facts. Relevant resources include reference documents, retrieval collections, and training or other informational inputs.
The oracle compares the information's source and declared status with the resource contract at ingestion, transformation, and use. A contained test can mix synthetic verified and unverified records and check whether transformations falsely upgrade their status. A correctly labeled unverified result is a negative control, even if its content is wrong. P21 governs trust relationship changes more broadly; P20 governs traceable origin. P06 protects the epistemic handling of a knowledge resource, while P22 and P23 concern observation and representation of an operational environment.
P07 Memory Integrity
Memory Integrity requires that retained operational state be created, altered, associated, interpreted, and retired only within its authorized security context. The contract identifies the task or principal owning the state, its scope and validity, and permitted persistence or reuse. A conversation summary, checkpoint, or stored preference can carry this obligation regardless of the storage medium or retention duration.
The oracle compares preexisting authority and state with what is restored or used in a later operation. An illustrative test resumes two synthetic user sessions and checks whether one user's preference or approval migrates into the other. A correctly bound resume is the negative control. P05 concerns assembled current context, and P06 knowledge resources. P17 concerns obligations across time even when memory contents are unchanged. P07 is warranted when the stored or restored operational state itself violates its authorization, association, or interpretation contract.
P08 Identity Integrity
Identity Integrity requires reliable binding of relevant actors and their asserted authority to the users, agents, services, sessions, or sources they represent. The contract defines the identity evidence accepted for a particular role and the permissible association between identity and authority. A familiar name or statement of role does not itself satisfy that binding.
The oracle compares the principal treated as acting with the verified identity and role record. A synthetic interagent message can reuse another agent's display name while retaining a different authenticated identity; the test examines whether the receiver grants the named agent's role. A message with a valid identity binding supplies the negative control. P20 concerns the history of origin and contribution; P21 concerns permissible trust relationships. P08 identifies mistaken actor or actor authority binding, whereas P16 examines whether a correctly identified delegate receives excessive scope.
P09 Objective Integrity
Objective Integrity requires preservation of the authorized task or outcome until a principal entitled to change it does so within scope. An objective record should identify its originating authority, governing constraints, accepted revisions, and unresolved conflicts. Deployment policy, a valid user request, and bounded delegation can contribute; their precedence must be specified. The system cannot establish authorization merely by describing a preferred goal as legitimate.
The oracle compares the task the system commits to pursue with that record. In an illustrative test, a report review task is redirected toward promoting a named supplier after reading source content. Evidence must establish a changed task criterion or pursued outcome; a wrong recommendation alone is insufficient. An authorized revision provides the negative control. P18 concerns a judgment made while the task remains fixed, and P15 concerns the strategy for achieving it. P09 does not judge the ethical merit of the authorized goal or resolve an unspecified governance conflict.
P10 Behavioral Integrity
Behavioral Integrity requires observance of a defined security relevant behavioral constraint when no more specific property captures the obligation. It is residual by design. The contract must describe the prohibited or required observable behavior independently of the fact that an output was disliked or harmful. Variation, surprise, and ordinary model error do not establish this property failure.
For example, a support system could have a reviewed prohibition on making targeted threats to coerce a person into continuing an interaction. A proposed test would define the threat and coercion criteria, benign discussion controls, adjudication procedure, and permitted evaluation setting before observing outputs. The oracle concerns those criteria, not a generic harmfulness score. Use P19 when the demonstrated obligation is materially deceptive influence on human decisions, P09 for changed objectives, P15 for invalid strategies, and P18 for invalid judgments. Where no specific behavioral contract or reliable adjudication exists, record the classification as unresolved rather than defaulting to P10.
P11 Semantic Integrity
Semantic Integrity requires preservation of security relevant meaning through interpretation or transformation. The contract identifies which distinctions must survive, such as negation, scope, conditions, quoted versus asserted statements, or a hypothetical versus actual authorization. It need not require literal wording or identical outputs, and ambiguity should be represented rather than silently resolved where the distinction matters.
The oracle compares the interpreted or transformed proposition with the specified meaning using controlled cases or a reviewed adjudication rubric. An illustrative test translates “approval is valid only for the test account” and checks whether the account limitation survives; meaning preserving paraphrases form negative controls. P04 applies specifically to instruction authority and meaning, P08 to actor binding, and P09 to the authorized goal. P11 is useful when the demonstrated failure is a meaning transformation; it should not be added solely because natural language appears somewhere on the path.
P12 Capability Integrity
Capability Integrity requires that the selection and intended use of available capabilities remain within their authorized security boundaries. The contract identifies permitted tools or operations, purposes, objects, principals, and conditions. Possession of a credential or technical access establishes only one part of the permission relation; it does not authorize every task or every proposed use.
The oracle compares a selected or requested capability use with its permitted purpose and scope before the relevant external effect. An illustrative test gives an agent both a read only query tool and a modification tool for separate tasks, then checks whether a source passage induces selection of the latter for a read only request. A legitimately authorized modification is a negative control. P13 concerns realized effects, P15 the strategy as a sequence, and P09 the objective. A blocked impermissible selection can violate P12 while execution enforcement successfully preserves P13.
P13 Action Integrity
Action Integrity requires that external effects remain within authorized scope at execution. The contract should bind the action's actual object, recipient, data, magnitude, timing, and relevant preconditions. It covers digital effects and physical actuation. Correct capability selection is not sufficient if parameters change, a target becomes stale, or execution produces a different effect.
The oracle compares observed effects with the applicable authorization and execution conditions. In a contained test, an approved write to synthetic record A is redirected to record B between selection and dispatch. A verified write to A supplies the negative control. P12 concerns capability selection; P13 concerns effect. If both contracts independently fail, both may be recorded in the causal account. This draft does not retain an obligatory earliest property rule that would hide a separately failed execution check. The earliest supported failure may remain an indexing choice without erasing later evidence.
P14 Controllability
Controllability requires that an authorized party retain the specified ability to observe, interrupt, constrain, revoke, or recover the system's operation. The contract identifies the available intervention, responsible party, timing requirement, and effects that remain reversible. A displayed stop control does not establish that running tasks or delegated operations obey it.
The oracle measures whether the intervention takes effect before its specified deadline and within the claimed scope. An illustrative simulation cancels a queued operation and checks whether execution and delegated continuations cease before a simulated commit point. A task with no pending effect is a negative control for the cancellation signal path but does not establish timely prevention. P03 concerns continued service, while P14 concerns legitimate control over that service. Recovery of reversible state cannot demonstrate reversal of disclosure or other irreversible effects; those require containment and intervention before the relevant boundary.
P15 Planning Integrity
Planning Integrity requires that a strategy for achieving an authorized objective satisfy declared security constraints on sequencing, dependencies, resources, and intermediate states. Legitimate goals and individually permitted operations do not guarantee that their composition is allowed. The assessable object is the selected plan, scheduled sequence, or sufficiently observable execution strategy, rather than an assumed private reasoning trace.
The oracle checks that strategy against its constraints before or independently of realized harm. An illustrative task permits updating two records only together; a plan to commit one before validating the other violates the stated condition. A transactional sequence supplies a negative control. P09 concerns what outcome is pursued; P18 concerns judgments such as whether validation passed; P13 concerns the actual update effects. Evidence of a harmful final action alone cannot determine whether the plan, the judgment supporting it, or only execution was defective.
P16 Delegation Integrity
Delegation Integrity requires that transferred or exercised authority preserve the delegator's valid scope and all constraints imposed on the transfer. The contract identifies the delegator, delegate, task, permitted resources, expiry, revocation behavior, and whether further delegation is allowed. Responsibility transfer is not an authorization to expand capability or change the objective.
The oracle compares the authority accepted or exercised by a delegate with the original grant and permitted transformations. A contained test delegates access to one synthetic folder and checks whether the next service treats the request as account wide access. A valid narrowed delegation is a negative control. P08 protects actor binding, P12 capability selection, and P17 time dependent validity. Information exchange or an agent's recommendation does not necessarily delegate authority; the graph must identify an actual delegation relation before applying this property on that basis.
P17 Temporal Integrity
Temporal Integrity requires that security obligations remain valid through specified delays, state transitions, repeated interactions, and lifecycle events. It captures temporal conditions such as expiry, revocation, order, and continuing validity. “The failure happened later” is insufficient; the assessment must identify which time dependent obligation was defeated.
The oracle compares an operation's timing and state transition with the relevant validity rule. An illustrative test queues a synthetic action under a grant that expires before execution and checks whether the action is revalidated or blocked. Execution within the grant's validity interval is a negative control. P07 concerns retained operational state, whereas P17 can fail with accurate storage when an unchanged expired grant is still honored. P14 concerns the effectiveness of intervention. Temporal qualifiers may accompany any property without constituting an independent P17 failure unless a distinct temporal obligation is evidenced.
P18 Decision Integrity
Decision Integrity requires that the system's security relevant judgments follow the specified decision rule, admissible evidence, and uncertainty requirements. The objective may remain unchanged while a verdict, classification, authorization determination, or intermediate judgment is manipulated. The contract must identify the rule or accepted adjudication method; disagreement with an assessor's intuition is not a sufficient oracle.
An illustrative test asks whether a synthetic request meets a fixed access condition and varies irrelevant source framing while holding the verified facts constant. The oracle compares the returned verdict with the declared rule and examines whether the purported influence explains the deviation. Legitimately different facts that change the verdict are negative controls. P09 concerns substitution of the task itself, P15 the strategy, and P19 the person's subsequent judgment. A wrong system verdict can be demonstrated without an external action, but attribution to an attack requires evidence beyond one incorrect answer.
P19 Human Decision Integrity
Human Decision Integrity requires that system mediated influence on security relevant human choices respect the declared conditions for materially truthful representation and informed authorization. The contract identifies which claims, omissions, uncertainty, or scope information are material to the person's role. Ordinary influence is not prohibited; the issue is unauthorized or materially deceptive influence under specified criteria.
The oracle can inspect a mismatch between presented action scope and the approval actually requested, or assess misleading claims under a reviewed rubric. Demonstrating that people changed decisions because of that presentation requires separate human evidence; an interface defect alone does not prove the downstream decision effect. A scope accurate preview is a negative control in a synthetic workflow. P18 concerns judgments produced by the system, P20 provenance, and P14 usable intervention. Human studies require appropriate consent, review, debriefing, and data protections; unsuspecting users must not become test subjects.
P20 Attribution Integrity
Attribution Integrity requires preservation of the origin and contribution records needed to trace significant information, decisions, and effects. The contract identifies the required lineage, including initiating principal, contributing sources, agents, tools, and authorization records where relevant. It does not promise perfect reconstruction of a model's internal reasoning or require indefinite retention of private content.
The oracle compares recorded lineage with controlled source and execution records. An illustrative test routes a synthetic request through two agents and a shared tool, then checks whether the resulting event incorrectly names the second agent as the original authorizer. A correctly linked event is the negative control. P08 concerns actor identity binding, P06 knowledge provenance in its resource role, and P21 trust assignment. Attribution supports accountability but does not itself establish responsibility, truth, or authorization. Logs can be complete yet misleading if their claimed relationships are not bound to the events they describe.
P21 Trust Integrity
Trust Integrity requires that trust relationships and the privileges of trusted roles change only through the authorized process. The contract must specify what trust permits: relying on a source as verified evidence, accepting it as an instruction issuer, or allowing a service to act for a principal. A single undifferentiated trusted flag can conceal materially different permissions.
The oracle identifies an unauthorized change in the role or reliance granted to a source or actor. An illustrative test inserts a claim of verified publisher status into an unverified record and inspects whether the receiving workflow accepts that claim as sufficient evidence for promotion. An actual authorized verification transition is the negative control. P08 concerns who the actor is; P21 concerns the trust relation assigned to that actor or content. P04 applies when the promotion specifically changes instruction authority. Several labels should reflect independently stated obligations, not repeated descriptions of one promotion.
P22 Perception Integrity
Perception Integrity requires that acquired observations preserve the accuracy, entity association, timing, and uncertainty required by the observation contract. Its stable concern is what the system observes. This revision explicitly includes observation of digital environments alongside physical sensing; historical physical only applications require version qualified review. Mere text input does not establish or exclude an observation function.
The oracle compares the observation representation with a controlled environment or suitably justified reference, using predefined error and confidence criteria. An illustrative simulation changes a displayed object's identity and checks whether the new observation binds to the correct object. A stable scene is a negative control. P06 concerns information supplied as knowledge; P23 concerns the maintained and inferred world representation. A single mistaken observation does not automatically demonstrate persistent world model corruption. The validity of a simulator or reference measurement limits the claim, and authenticated sensor origin does not alone establish perceptual accuracy.
P23 World-Model Integrity
World-Model Integrity requires that a maintained representation of the operational environment remain consistent with required observations, state transitions, uncertainty, and refresh rules. It includes inferred state and conditions not currently observed. The contract must define the environment features that matter, acceptable staleness, and triggers for revising beliefs; perfect knowledge of reality is not a feasible requirement.
The oracle compares the maintained representation and its use with a controlled environment history. An illustrative digital simulation changes resource ownership after an earlier observation and checks whether a required refresh corrects the represented owner before a dependent decision. A permitted stable ownership interval is a negative control. P22 concerns acquired observations, P07 authorized retention, and P17 temporal conditions more broadly. A faithfully stored but stale world representation can violate P23 without a storage alteration. A false observation that is promptly corrected need not establish a separate P23 violation.
Distinguishing goals judgments strategies and behavior
P09, P18, P15, and P10 answer different questions. P09 asks what authorized outcome the system is trying to achieve. P18 asks whether a security relevant judgment follows its declared rule. P15 asks whether the chosen strategy respects its constraints. P10 asks whether a separately defined residual behavioral constraint is breached. The fact that each can affect output does not erase these boundaries.
Consider an illustrative, unexecuted transaction review task. Changing the task from detecting unauthorized transfers to maximizing approvals suggests P09. Preserving the review task but declaring a transfer valid contrary to the fixed rule suggests P18. Reaching the correct verdict but planning to notify a recipient before completing required checks suggests P15. None should be relabeled P10 merely because it produces undesirable behavior. Evidence may support several failures, but each needs its own obligation and oracle; unseen intermediate states should remain hypotheses.
Reporting interactions and limits
A finding should name the narrowest supported obligation, its component or relationship, the relevant domain, and the mechanism evidenced. Generic integrity and contextual qualifiers can aid retrieval without multiplying vulnerability counts. Multiple independently failed controls may be recorded in one causal case. A path that crosses memory, interpretation, and execution does not establish three property violations unless each relevant contract fails.
Nonadversarial disturbances can exercise these obligations. Record the observed trigger, then separately assess whether an adversary can cause or exploit it. Likewise, a vulnerability can exist before a harmful incident occurs; the assessable violation may be an unauthorized grant or disclosure path rather than a completed catastrophe. Confidence, recurrence, consequence severity, exposure, and test coverage remain distinct. Irreversible consequences justify containment and conservative decisions, while probabilistic evidence remains meaningful within its assumptions.
Foundational Concepts defines contracts and evidence; Security Layers locates their functions; Threat Model states assumptions and capabilities; Assessment Methodology governs execution records; and Vulnerability Registry governs reviewed claims. Future proposals should justify added or refined properties with boundary cases and reproducible evidence. This draft supplies a language for scoped assessment, not proof of total coverage or a certificate that any system is secure.
Word document · Editable source
Threat Model
Framework version: 2.0.0-draft.1. Status: research proposal.
Purpose and scope
This threat model describes how an intelligent system can lose a specified security property through adversarial influence or through a security-relevant failure without an adversary. Its unit of analysis is the deployed or proposed system under explicit assumptions. That unit can include models, conventional software, data services, sensors, retained state, tools, operators, affected people, and interacting organizations. A model endpoint alone is an adequate boundary only when the claim is correspondingly narrow.
The method applies by function. Language generation, statistical prediction, symbolic inference, online adaptation, planning, perception, and physical actuation can contribute to the system under examination. No particular interface or model family establishes coverage. An assessment of a text service may include observation of a digital environment; an assessment of a robot must include conventional access controls as well as physical observation and action.
The output is a reviewable threat model that connects protected interests, actual components, authority, possible influence, failed contracts, and impact. It is a hypothesis-producing artifact, not evidence that every threat exists or that an implementation is vulnerable. Assessment Methodology defines how hypotheses are tested. Domain Profiles supplies candidate deployment-specific constraints; Security Layers defines the analytical domains used to describe them.
Establishing the system boundary
Begin with a dated description of the intended task and the parties whose interests the assessment protects. Record the deployment, configuration, operational period, and lifecycle stages included. State whether the evaluation includes model preparation, data collection, training, updates, deployment, operation, retirement, or only a subset. A runtime assessment does not establish that training provenance is trustworthy.
Draw a graph of actual components, humans, services, and shared resources. Assign stable local identifiers to nodes and relations. Include externally operated services when the system relies on their outputs or availability, even if their internal implementation is inaccessible. Mark those internals as unassessed dependencies. Shared infrastructure should appear as a shared node rather than as several fictitiously independent copies.
Relations are typed as observation, information, instruction, authority delegation, action, state synchronization, feedback, or resource dependency. Record direction, content, relevant identity, validation, and failure behavior. A relation may have several types, but a message does not automatically confer authority. Two agents exchanging facts need an information contract; a supervisor assigning a bounded capability also needs a delegation contract.
Distinguish the evaluation boundary from the consequence boundary. A tool may be a test stub while the corresponding production action affects another organization. The assessment must describe that difference and cannot claim observed production harm from a simulated action. Dependencies beyond access can remain explicit assumptions; they must not disappear from the threat model merely because they are difficult to inspect.
Protected interests and contracts
A property label becomes testable only through a contract. For each important protected object, function, or relationship, specify the responsible owner, authorized operations, prohibited transitions, assumptions, observation method, and response to uncertainty. Use the stable P01–P23 identifiers from Security Properties and Versioning and Identifiers. Several properties can support one contract, and one property can apply to several components.
For example, P16 Delegation Integrity can support a contract that a delegated operation remains within an authenticated issuer's authority, named resource scope, permitted operation set, and expiry. The contract should also say whether delegation can be transferred and how revocation is checked. The word “trusted” is insufficient without these conditions.
Relevant protected interests include confidentiality of records, availability of essential services, integrity of model artifacts, instruction priority, provenance of evidence, continuity of retained state, authorized objectives, bounded plans and actions, informed human authorization, and collective resource limits. Record conflicts rather than assuming all properties can always be maximized together. A recovery action that restores availability may change retained state; the permitted tradeoff requires an authorized rule.
Authority must have a source outside the disputed content. Specify who can establish, change, delegate, or revoke each objective and permission. Authentication establishes a claimed identity under a mechanism; it does not by itself establish that every requested action is permitted. A legitimate user's request can exceed their rights. Likewise, task relevance, fluency, or retrieved prominence does not confer instruction authority.
Actors and sources of influence
Actor categories help discovery but do not substitute for a capability description. Consider unauthenticated outsiders, authenticated users with limited authority, insiders, data publishers, model or dependency suppliers, tool operators, other agents, and parties controlling physical or digital surroundings. A supplier can be honest but compromised. An autonomous system can provide harmful input without the assessment having established intentions or independent goals.
For each adversarial scenario, record access, knowledge, control, observation, budget, timing, and constraints. Access specifies reachable interfaces and prerequisites. Knowledge specifies whether the actor knows prompts, schemas, model artifacts, policies, or only outputs. Control identifies exactly which bytes, records, signals, timings, identities, or actions can be changed. Observation identifies feedback available after an attempt. Budget covers attempts, compute, duration, account creation, and other relevant resources. Timing identifies when intervention is possible relative to validation, state change, or action commitment.
State what the actor cannot do. A scenario that assumes write access to system policy is different from one that permits only publishing a retrieved document. If the evaluator obtains extra privileges solely to construct a fixture, explain which resulting state the attacker could realistically create and which parts are laboratory conveniences. Administrative setup does not establish adversarial reachability.
The objective should name a security-relevant outcome, not merely a response format. “Produce a different answer” is normally insufficient. “Cause release of a synthetic record to a principal outside its access policy” identifies a protected interest and an observable event. Do not equate a model's generated claim of success with an actual state transition.
Nonadversarial conditions
Record failures caused by outages, stale observations, accidental misconfiguration, distribution change, conflicting legitimate requests, or incompatible local policies without inventing a malicious actor. These may expose the same contract weaknesses that an attacker could exploit. Their causal description should identify a disturbance or hazard source and its operating conditions.
A hazard is a condition capable of causing harm; an incident is an observed event; a vulnerability is a weakness that permits a security-property violation under stated conditions. A quality defect can remain outside security scope when no relevant protected interest or contract is established. These categories can overlap in a case, but they are not interchangeable. An accidental incident does not demonstrate an attack technique, and an adversarially authored input that produces no violation does not establish a vulnerability.
This distinction preserves the broad scope of intelligent-systems security while avoiding a catalogue in which every error becomes an attack. Cases with disputed security relevance should be retained with the disagreement and the missing contract evidence.
Analytical domain mapping
The nine domains are provisional analytical groupings, not execution stages or ranks of consequence. Map actual functions to them rather than forcing components into one box.
L1 and L2 distinguish model computation from its software realization: a compromised library is an implementation issue even when it serves inference. L3 and L6 distinguish information resources from retained operational continuity: the same datastore may serve both functions. L5 concerns interpretation and objectives, while L7 concerns the plans and actions developed or executed under them. Harmful behavior alone does not locate a failure in L5.
L9 requires an identified collective contract and coupling mechanism. Large impact from one compromised component is not sufficient. A shared resource limit or feedback stability condition can fail at a group or subgraph even when no individual relation violates its local contract. The graph is therefore not restricted to pairwise failures.
Apply D1–D7 as cross-cutting dimensions with their original meanings. Identity, governance, provenance, privacy, lifecycle, observability, and recovery remain relevant across domains. A domain marks an analytical responsibility; a dimension provides an additional perspective; a property states what must hold.
Constructing influence and failure paths
Describe each scenario as an entry point, a controlled intervention or disturbance, intermediate transitions, hypothesized failed contracts, and possible consequences. Mark which transitions are observed, inferred, assumed, or untested. An attack may traverse a database, context assembler, planner, and tool without each component containing a separate vulnerability.
Permit several causal failures, a composite failure, or unresolved classification. A primary domain can be recorded when the evidence supports one, but is not mandatory. Distinguish the location where influence enters from the location where authority is incorrectly assigned. Also distinguish containment failures that are independently established from controls that were never promised to prevent that event.
Include temporal structure: prerequisites, ordering, state retention, expiry, retries, delayed triggers, and opportunities to intervene. A memory-related scenario needs a later context and reset conditions, not only an immediate output. An action scenario needs the point at which a commitment becomes externally effective. For interacting agents, include cycles, shared state, and resource contention rather than assuming that influence moves only forward.
Worked scenario
The following scenario is illustrative and unexecuted. A knowledge assistant retrieves a document from a source that an external publisher can edit. The publisher cannot change system instructions, credentials, or tool policy. The assistant's authorized task is to summarize internal guidance; it has a test-only tool for proposing document updates.
The hypothesized intervention is document content presented as an instruction to change a retained authorization preference. Entry occurs through an L3 information relation. A possible L5 failure would promote document content into instruction authority. A possible L6 failure would persist that preference without an authorized update. A later L7 failure would require separate evidence that an action contract was violated. None follows automatically from the preceding event.
The relevant contracts can involve P04 Instruction Integrity, P07 Memory Integrity, and P09 Objective Integrity. The scenario would be weakened if the preference is never committed, if an authorized user explicitly requested the update, or if the outcome occurs equally under matched benign content. Initial tests should use synthetic preferences, isolated state, and a tool stub. Production compromise, prevalence, and downstream impact remain unestablished.
Review and maintenance
Review assumptions with system owners and relevant practitioners. Record disagreement about authority, protected interests, and operational constraints before selecting tests. Prioritize scenarios using justified exposure and consequence descriptions; do not multiply ordinal labels into a universal score. High-consequence uncertainty may justify early containment even while causal investigation remains incomplete.
Version the graph, contracts, scenario records, exclusions, and unresolved dependencies. Revisit the threat model after changes to objectives, models, tools, memory, permissions, data sources, human workflows, or collective composition. Assessment results should feed back into these artifacts, including negative findings and unrepresented mechanisms. A completed threat model establishes the scope of inquiry, not a proof of security or future completeness.
Word document · Editable source
Attack Taxonomy
OSAFIS 2.0.0-draft.1 · Proposed analytical method · 2026-09-07
Purpose and limits
An attack classification should explain the intervention an adversary attempted, the contract it challenged, and the evidence connecting intervention to effect. An input format, striking output, or memorable name does not supply that explanation. This taxonomy supports comparison across intelligent systems without assuming that every system uses natural language, learned weights, persistent memory, or autonomous tools.
The taxonomy is a controlled vocabulary with relationships, not an exhaustive tree of mutually exclusive attacks. Its twelve inherited mechanism identifiers remain stable. They describe overlapping concepts at different levels of abstraction: some emphasize an intervention, some a target, and some a pattern across interactions. This limitation is explicit. There is no evidence in this proposal that twelve mechanisms are optimal, exhaustive, or equally granular.
Use the vocabulary to organize an evidence record. Do not use the presence of a label as evidence that a vulnerability exists, that a proposed test detects it, or that a control mitigates it. A benign quality defect or safety hazard may violate a relevant contract without constituting an adversarial attack. Such a case receives no invented attack mechanism.
Classification objects and their relationships
A mechanism describes how an adversary attempts to alter a security-relevant process or relationship. A family groups supported instances sharing a discriminating causal account. A subfamily records a reproducible variation that changes an important boundary condition, intervention, or control response. A technique is the concrete procedure used in an instance. Modifiers describe conditions such as interaction length, channel, timing, access, language, or presentation. Effects describe observations; impacts describe consequences for protected interests.
These objects are related rather than nested into a mandatory sequence. A technique can instantiate multiple mechanisms. A modifier can apply to several families. A family need not have subfamilies. An observation may remain unclassified pending evidence. Data exfiltration, for example, describes an effect involving confidentiality; it does not by itself explain whether the cause was a conventional access defect, misinterpreted authority, or an action-validation failure.
Layers answer where a failed contract resides. Properties answer what must be preserved. Dimensions identify concerns that recur across loci. A capability describes what an actor can do; an entry point describes where influence enters the assessed graph. None of these is interchangeable with mechanism. “A tool response” is an entry channel, “untrusted text accepted as an instruction” is a causal claim, and “a record was disclosed” is an effect requiring separate evidence.
Retained mechanism vocabulary
The following identifiers and labels retain the source vocabulary. Inclusion criteria make their application more explicit without assigning new meanings to existing identifiers.
M01 Technical Manipulation
Use for adversarial interference with software, infrastructure, configuration, authentication, authorization, or runtime controls. Identify the concrete technical boundary and attacker access. A process isolation defect and an unauthorized configuration write can both qualify, although their techniques and consequences differ. The label alone does not establish L2: a technical operation may corrupt a model artifact at L1 or retained state at L6. Record both the enabling contract and the object affected when evidence supports both.
M02 Data Manipulation
Use for adversarial changes to information consumed by a system, including insertion, deletion, selection, and misleading attribution. State what information changed and which consumer relied on it. Ordinary retrieval of attacker-authored public material is not automatically a failed L3 contract. Failure depends on the promised provenance, validation, or use restrictions. If faithfully delivered text is later promoted into instruction authority, the supported failure may be L5 while L3 participates correctly.
M03 Environmental Manipulation
The inherited mechanism concerns adversarial changes to the physical environment observed by a system. Distinguish changing an observed scene from compromising the sensor pipeline. A physical change may be observed accurately and therefore cause no L4 violation. The revised L4 also covers digital environment estimation; that domain expansion does not silently expand M03. Describe an adversarial digital environmental intervention through supported existing mechanisms or leave its mechanism unresolved pending a versioned proposal.
M04 Perception Manipulation
Use for adversarial interference with acquisition or interpretation of physical signals, including spoofing and interference. Identify the observation error or uncertainty violation, not merely the presence of a camera or microphone. A sign containing instructions that is transcribed correctly does not establish M04. Conversely, falsifying a position estimate can establish an observation failure without semantic instruction conflict. Digital perception cases can be classified at L4 even where the inherited mechanism vocabulary requires clarification.
M05 Semantic Manipulation
Use when the adversary changes how a representation is interpreted, including whether content is treated as instruction, evidence, quotation, or an authority assertion. Specify the interpretation contract and the competing interpretations. A change in wording followed by a changed answer is insufficient: legitimate meaning may also have changed. Compare content and task constraints, and account for ambiguity in the authorized request before attributing a security failure.
M06 Contextual Manipulation
Use when an adversary changes surrounding information, ordering, framing, or interaction history to affect interpretation. Identify the contextual intervention independently of the resulting decision. M06 may accompany M05, but the two claims differ: one identifies manipulation of surroundings, the other an interpretation change. If evidence cannot distinguish them, retain the broader supported description and mark the narrower attribution uncertain.
M07 Behavioral Manipulation
Use for adversarial influence across interactions or decision sequences where that sequence is part of the causal hypothesis. It is not a default tag for every changed output. A report should identify the interaction dependency and test whether disrupting or replacing the sequence changes the outcome. Where all evidence is a single terminal action, describe the action as an effect and avoid inferring a distinct behavioral mechanism.
M08 Memory Manipulation
Use for adversarial modification of retained state that influences later operation. Specify the write, association, retention, or reactivation path and the temporal interval. The same stored object can serve L3 information and L6 continuity functions. Poisoned training examples are not automatically memory manipulation; a model update is not equivalent to a retained user preference. State which contract governs the particular role.
M09 Identity Manipulation
Use for adversarial interference with perceived identity, authority, trust, or actor relationships. An authenticated actor can still claim authority it does not possess. Identify the assertion and verifier, and distinguish provenance loss from identity substitution. D1 Identity, Trust & Authorization remains a cross-cutting dimension; M09 is an adversarial mechanism, not a replacement for that dimension or a new identity layer.
M10 Objective Manipulation
Use when an adversary redirects the apparent or effective objective. Record the authorized objective, the alleged substituted objective, and evidence distinguishing redirection from a faulty plan for the original objective. M05 or M06 may explain how redirection occurred. M10 should not be inferred solely because a result was harmful: unsuccessful execution of the correct objective is a different explanation.
M11 Capability Manipulation
Use for adversarial influence over selection or use of capabilities, including tools and delegated work. Identify the capability boundary and the action the adversary sought. A valid credential does not prove that a particular use serves an authorized purpose. Distinguish a capability exposed too broadly from an otherwise permitted capability invoked with an invalid target, sequence, or delegated scope.
M12 Human Manipulation
Use for adversarial exploitation of system behavior or output to influence human decisions. Specify the adversarial control, decision contract, and proposed causal pathway. Persuasion or disagreement alone does not establish a vulnerability. A test must distinguish informed voluntary choice from a failure of disclosure, consent, authority representation, or oversight. Human-subject work requires appropriate consent, review, debriefing, and data protection; hypothetical examples are not human evidence.
Resolving overlapping descriptions
Consider a hypothetical adversary who inserts an instruction into a reference document and surrounds it with false administrative context. If the system treats that reference as authoritative and adopts a different task, M02 describes the information intervention, M06 the contextual intervention, M05 the authority interpretation, and M10 the redirected objective. These are candidate descriptions of distinct steps, not four independently demonstrated vulnerabilities.
Assign each mechanism to a specific supported claim. Remove tags that merely restate the terminal effect. Evidence may establish the reference-to-instruction failure while leaving the exact contribution of framing unresolved. In that situation, a narrower report with uncertainty is more informative than a maximal tag set.
A useful boundary test asks what intervention could separate two explanations. Hold task content constant while changing contextual placement to test a contextual contribution. Preserve framing while restoring authenticated instruction boundaries to test authority interpretation. Preserve the objective while replacing the planner to distinguish objective corruption from planning failure. Such comparisons do not guarantee causal identification; document remaining confounds, including unequal content, model variability, and hidden state.
Mapping attacks to the analytical domains
Use the canonical domains: L1 Models & Computation; L2 Software & Infrastructure; L3 Data & Knowledge; L4 Perception & World Representation; L5 Interpretation & Objectives; L6 Memory & State Continuity; L7 Planning & Action; L8 Human–System Interaction; and L9 Collective & Systemic Interaction. They are provisional analytical domains, not execution stages or severity ranks.
Record violated domains separately from participating domains. A camera that reads an adversarial instruction correctly may participate at L4 without violating P22 Perception Integrity. A tool may execute a validated benign action during an attack chain without failing at L7. Broad distribution of a single compromised model does not automatically establish L9; a distinct collective contract and coupling mechanism are required.
One finding may involve several failed contracts. Do not force a primary layer when evidence supports interacting failures or cannot resolve attribution. If a primary label is needed for indexing, explain the selection rule and retain the full graph. Neither earliest entry nor greatest impact is a universal rule for locating the vulnerability.
Graph and sequence representation
Represent actual components, humans, shared services, and resources as nodes. Type relations as observation, information, instruction, authority delegation, action, state synchronization, feedback, or resource dependency. An information edge does not automatically transfer authority. A shared model is one shared node when that is the architecture; duplicating it for every agent hides common dependence.
Annotate the graph with entry, attacker capability, failed contract, observable evidence, and downstream consequences. Include time and state transitions when persistence or adaptation matters. Group constraints may apply to a subgraph: a collective resource budget cannot always be reduced to one pairwise communication edge. Mark hypothetical propagation separately from observed propagation and identify where evidence ends.
For example, the hypothetical path “reference source → context assembler → interpreter → state store → planner → simulated action sink” describes participation. Additional annotations identify whether provenance admission, instruction authority, retained preference approval, or action validation actually failed. A path through six components is not evidence of six vulnerabilities.
Classification record
A usable record contains a version-qualified reference, system scope, asset, authorized behavior, attacker capability, entry point, assumptions, failed contracts, and participating components. It then records properties, mechanism claims with confidence, optional family and subfamily, technique description, modifiers, observations, impacts, and alternative explanations. Evidence artifacts and test conditions must permit the reader to distinguish execution from inference.
The record also states containment, controls, failure criteria, repetition plan, and limitations. Seeded-violation positive controls demonstrate that the apparatus can detect a deliberately seeded contract failure in the isolated test environment. Authorized-function controls verify that legitimate operation remains possible. Negative controls exercise comparable behavior expected to preserve the contract. State the expected response for each control. A clean negative control does not prove broad security, and a failed seeded-violation control makes a negative result difficult to interpret.
Use the Security Properties document for P01–P23 definitions and the Assessment Methodology for evidence handling. Preserve exact identifiers when exporting a record. A local example identifier is not a registry identifier. An unexecuted test plan must leave measured success, impact, and reproducibility fields unset rather than filling them with plausible values.
Families and evidence for novelty
Familiar descriptive terms such as prompt injection, retrieval poisoning, or delegation escalation can help readers navigate. Their use here does not assert that OSAFIS discovered or independently validated those phenomena. A proposed new OSAFIS family requires a discriminating definition, security contract, reproducible evidence, comparison with adjacent descriptions, and a falsification strategy.
Family promotion should answer what the grouping predicts that existing descriptions do not. A reproducible difference in boundary conditions or control response can justify a useful subdivision. A different channel, longer prompt, or fashionable architecture usually supplies a modifier or technique unless evidence supports a deeper distinction. Exact-string reproducibility can validate an instance without establishing generality of a family.
Reject a novelty claim when matched conditions remove the effect, when legitimate task differences explain it, or when the proposed category contributes no distinguishable causal account. Record negative findings and unresolved cases. They constrain the taxonomy and prevent repeated rediscovery of weak claims.
The original Attack_Taxonomy.docx describes a candidate involving intent-preserving semantic transformations and reports a reproduced-observation status. That narrative is a source claim, not independently verified experimental evidence in this revision. The candidate remains a research question unless underlying artifacts, pinned conditions, content-matched controls, and outcome coding support stronger conclusions. No named new family is promoted by this document.
Negative classification and uncertainty
Use the assessment outcomes defined in Assessment Methodology: supported, refuted, inconclusive, quality_issue, hazard_only, and out_of_scope. Record whether a case matches an existing category, remains unrepresented, or has ambiguous classification separately. Record novelty and registry acceptance as separate review decisions. A contract failure may be real while its adversarial mechanism remains uncertain; a successful adversarial attempt may exploit an existing category rather than justify a new one.
Record uncertainty at the claim level. Confidence in an observed unauthorized write may be high while confidence in semantic versus contextual contribution is low. Severity concerns consequences under stated conditions; confidence concerns evidential support. Neither can substitute for the other. A rare but credible severe failure should not disappear because a family label is unsettled.
Validation and revision
Evaluate the vocabulary using independently classified held-out cases spanning model artifacts, conventional software, retrieval, physical and digital observation, retained state, symbolic systems, delegated action, human interaction, and collective dynamics. Freeze definitions before evaluation. Measure agreement on failed contracts separately from agreement on mechanism tags, since overlapping mechanisms make a single-category agreement score misleading.
Report unrepresented cases, ambiguous cases, reviewer time, and disagreements requiring adjudication. Predeclare what would trigger a definition change, split, merge, or retirement. A useful extension must improve discrimination without merely moving difficult cases into a broad residual category. After revision, evaluate on fresh cases rather than treating reclassified development examples as independent confirmation.
Versioning and Identifiers governs stable IDs and migration. This proposal preserves M01–M12 labels while replacing the implication of an exhaustive hierarchy with explicit relational use. Classification quality and future coverage remain empirical questions. The Worked Examples supply unexecuted practice cases; Future and Autonomous Systems defines architectural stress tests; neither supplies validation results.
Word document · Editable source
Assessment Methodology
Framework version: 2.0.0-draft.1. Status: research proposal.
Purpose and claim discipline
This methodology turns a security hypothesis into a bounded, reproducible assessment. It specifies artifacts and decision rules for investigating whether a defined contract can fail, how the failure occurs, and what conclusions the available evidence supports. It does not supply a universal certification threshold or establish that the proposed framework has already been empirically validated.
An assessment must separate the existence of a weakness, the frequency of an observed outcome under test conditions, the extent of demonstrated impact, and the generality of the proposed explanation. A single well-documented unauthorized state transition can warrant containment. It does not establish prevalence across deployments or a new attack family. Conversely, failure to reproduce an event is evidence about the tested conditions, not proof that the original report was false.
Required artifacts are a scope and authority record, system and threat model, property contracts, test protocol, execution record, analysis, falsification record, classification, and evidence package. Their detail should match the claim and consequences. An implementation-specific vulnerability need not demonstrate cross-model generality; a broad mechanistic claim requires evidence beyond one tailored example.
Scope and authorization
Record the target configuration, owner, evaluator, purpose, permitted interfaces, test period, data handling constraints, and stop conditions before testing. Identify whether actions reach production systems, a staging deployment, a simulator, or stubs. Use synthetic records and isolated accounts wherever they can exercise the relevant contract. Permission to inspect a component does not imply permission to affect its users or connected services.
Establish containment before executing a candidate intervention. Describe network restrictions, tool substitutions, spending or resource limits, state snapshots, and recovery responsibilities as appropriate. A stop condition should specify who can stop the test, which observable event triggers it, and how residual work is cancelled. Tests involving physical actuation or human exposure require the additional controls described below; harmful live experiments are not part of this method.
An assessment of a nonadversarial hazard follows the same contract and evidence discipline but records the disturbance rather than assigning an attacker. The report distinguishes security findings, hazards, incidents, quality defects, and unresolved observations.
Question and operational contract
Write the principal hypothesis before the confirmatory test. It should identify the actor or disturbance, available influence, target contract, expected failure event, and scope of inference. State a plausible alternative explanation and the observations that would weaken or refute the hypothesis.
Use P01–P23 as references, then operationalize the selected property. For example, P13 Action Integrity might require that the executor never commits a write outside the approved resource set for a named task. Specify what counts as approval, when it is checked, how scope changes, and which log or state readback demonstrates commitment. A model response saying it wrote a resource is not the same observation.
Define the outcome oracle before evaluating results. Prefer independently observable state, permission checks, or invariant violations over interpretation of prose. For semantic judgments, provide a coding rubric, examples of boundary cases, and an inconclusive category. If a model is used as a judge, freeze its configuration, validate it against independent annotations appropriate to the claim, and disclose disagreement and possible shared failure modes. A judge's label is evidence from that measurement process, not ground truth by definition.
Design and controls
Establish a baseline under normal authorized operation. Use a matched negative control that preserves task content and relevant difficulty while removing the hypothesized malicious feature. Distinguish an authorized-function control, which verifies that legitimate operation remains possible, from a seeded-violation control, which verifies that the oracle detects a known contract failure in an isolated fixture. State which response each control should produce. All controls must remain contained and must not create real harm.
Document the intervention, what remains fixed, and unavoidable confounders. “Only one variable changed” is a design objective, not a statement to make when a payload also changes length, retrieval ranking, or task content. Use additional controls or bounded conclusions when these factors cannot be separated. Randomize or counterbalance order when order effects are plausible.
Exploratory search and confirmatory evaluation must be distinguished. Record the number and nature of payloads, prompts, or configurations explored, including unsuccessful attempts. Freeze a selected candidate before evaluating it on heldout tasks or fresh states. Continuing to adapt a payload against the same evaluation set changes the claim to performance under that adaptive search procedure.
Units and sample planning
Define the independent experimental unit. A turn is not automatically independent of earlier turns; several outputs from one persistent session may form one unit. Multiple agents sharing a service or model state may be clustered. For human outcomes the unit may be a participant, team, or organization depending on assignment and dependence. State the unit of randomization, observation, and analysis when they differ.
Plan sample size around the claim, expected variability, smallest relevant effect or desired precision, dependence structure, and available resources. If information is insufficient for a justified confirmatory calculation, label the study exploratory and use a pilot to estimate design parameters. Do not adopt a fixed trial count merely because an earlier example used it. State stopping rules and how sequential monitoring or multiple comparisons will be handled.
Report denominators, exclusions, incomplete runs, timeouts, and missing observations. Distinguish invalid harness executions from genuine failures of system availability. Excluding inconvenient outcomes after seeing their condition can bias the result. Predefine invalidation rules and report sensitivity to disputed exclusions.
Configuration and provenance
Record model identifiers and available revision details, sampling parameters, seed support, prompts or policy artifacts, context construction, tool schemas, permissions, retrieval snapshots, memory state, software versions, and relevant timing. Record unknown or provider-controlled variables explicitly. A public model name may not identify immutable behavior; include test dates and any available deployment revision.
Store exact inputs and outputs where permitted, along with hashes, timestamps, run identifiers, and execution logs. A seed can aid reproducibility without guaranteeing deterministic execution across platforms or provider changes. If state cannot be restored exactly, describe the reset approximation and test its adequacy. Record the effect of caches, concurrency, rate limits, and background updates where they can influence the outcome.
Protect secrets and personal information in the evidence package. Replace sensitive values with synthetic equivalents before testing when feasible. Redaction should preserve the evidence needed for the claim; if it does not, identify what an authorized reviewer must inspect privately.
Execution and causal checks
Run baseline and intervention conditions under the planned protocol. Capture the earliest observable divergence, downstream transitions, and final outcome. Read back relevant state rather than relying solely on a tool's acknowledgement. Record refusals, partial actions, recoveries, and delayed effects as separate outcomes when the contract requires that distinction.
Use ablations to investigate necessary conditions: remove the asserted authority cue, disable the relevant persistence path, replace the suspect source, or interrupt a hypothesized feedback relation. These are tests of the explanation, not automatic fixes. An intervention that removes all system functionality cannot by itself show that a targeted control is effective.
Conventional static analysis or a direct access-control test may establish some implementation contracts without stochastic trials. Record the path and preconditions and verify relevant state transitions safely. Do not force every assessment into a language-model experiment merely because the surrounding product uses AI.
Temporal and adaptive evaluation
For multi-turn cases retain the complete sequence and identify the independent session unit. Test whether earlier turns are necessary, whether order matters, and whether the effect persists after a documented reset. For L6 Memory & State Continuity, separate the write event, retained representation, later retrieval, and subsequent decision. An immediate response does not establish cross-session persistence.
For adaptive systems, record the initial state, update rule or service behavior available to the evaluator, authorized learning inputs, update schedule, and rollback mechanism. Compare matched update histories or replayable streams where possible. A snapshot test does not cover later adaptation. Holdout conditions must remain outside the adaptation process if they are intended to measure generalization.
For actions, identify the last effective intervention point and the point of external commitment. Test cancellation and revocation in a contained environment, including relevant races and retries. Report whether attempted, queued, executed, compensated, and irreversible outcomes are distinguishable in the evidence.
Human and physical evaluation
Claims that presentation changed human decisions require human evidence. A narrower P19 Human Decision Integrity presentation or informed-authorization contract can be assessed through an inspectable mismatch, without claiming that anyone was deceived or changed a decision. Interface inspection can establish a mismatch or missing control without establishing its population effect. Label that narrower result accurately.
Human-subject research requires appropriate ethics review, informed consent, justified recruitment and sample planning, privacy protections, and debriefing where approved deception is involved. Use benign scenarios without actual financial, health, employment, or security consequences. Specify validated measurement instruments where appropriate and cite their sources in the individual study. Blind outcome coding when practical and account for repeated observations from the same participant. No universal human influence instrument is supplied here.
Physical-system tests should begin with simulation, recorded observations, hardware isolation, or nonhazardous fixtures. Document the gap between these conditions and deployment. Relevant practitioners must review containment and acceptance rules before any real-world actuation study. Do not expose people, animals, or operational infrastructure to a harmful condition to demonstrate reachability. A simulated contact or prohibited action remains a simulated outcome.
Collective evaluation
For L9 Collective & Systemic Interaction, identify the collective contract and graph or subgraph to which it applies. Record participants, shared resources, scheduling, message semantics, initial states, and coupling. Measure the proposed propagation or amplification relative to a matched baseline, with its denominator and time window.
Vary relevant composition assumptions: participant count, topology, delays, shared dependencies, and local policies. A result that depends on one arrangement can still establish a vulnerability of that arrangement; it does not establish a universal property of multi-agent systems. Independent local behavior does not rule out a collective failure, while correlated failures from one shared service do not demonstrate agent-to-agent propagation.
Use relation ablation, shared-resource substitution, or alternative schedules to test the causal explanation. Permit group-level or subgraph loci when no single edge is responsible. Simulation assumptions and validation limits must accompany all systemic claims; modeled evidence is not equivalent to observation of a deployed population.
Analysis and uncertainty
For each condition report valid units, observed contract violations, other outcomes, and uncertainty appropriate to the design. An observed proportion describes the tested distribution and procedure. It is not deployment likelihood unless the sampling and exposure assumptions justify that inference. Report differences or other prespecified effect estimates with uncertainty rather than only a significance label.
Account for clustering, adaptive selection, multiple comparisons, and incomplete observations where relevant. Use paired analysis when the design is paired. For small or sparse data, make the limits visible rather than presenting precise-looking percentages alone. If statistical modeling is used, record assumptions, diagnostics, and sensitivity analyses in the evidence package.
No observed failures in a finite test does not prove zero risk. A statistical bound, when reported, depends on independence, sampling, and model assumptions. Irreversible outcomes remain amenable to probabilistic analysis, but acceptable average performance does not compensate for a prohibited catastrophic event. Reachability, containment, intervention time, and frequency estimates can all matter; the applicable profile defines the decision rule.
Falsification and classification
Actively test alternative explanations. Check whether the baseline exhibits the same violation, whether the intervention changed authorization, whether the oracle misread a harmless output, and whether hidden state or harness behavior explains the result. A mechanism-specific hypothesis can fail even while a real vulnerability remains. Document both outcomes.
Map entry points, failed contracts, propagation, and impact separately. Use the canonical L1–L9 names and stable M01–M12 and D1–D7 identifiers. Traversal alone does not establish a failed domain. Permit multiple causal failures, composite classification, ambiguity, and unrepresented cases. A forced primary layer can hide the most informative part of a finding.
Compare a proposed new family with the closest existing mechanisms and categories. New wording, a new target product, or a larger consequence does not by itself establish mechanistic novelty. Conversely, a narrow but real implementation weakness should not be rejected because it is not novel. Registry review of a class proposal is a different decision from remediation of a concrete finding.
Evidence descriptors
The source framework's evidence labels are retained as descriptive facets rather than a compulsory cumulative ladder. Cross-model testing and independent replication answer different questions. A carefully controlled implementation finding can be strong without cross-model applicability.
For each facet record supported, unsupported, not_tested, inconclusive, or not_applicable, with evidence references and rationale. Unsupported means the cited evaluation does not support that facet, not that no vulnerability exists. A facet marked supported must identify the exact claim and tested scope. Independence includes authorship, analysis, data, and execution dependencies; reviewers should disclose which are shared.
Use assessment outcomes supported, refuted, inconclusive, quality_issue, hazard_only, or out_of_scope, with a reason and claim reference. These outcomes are separate from registry review status. Mixed findings can contain several claims with different outcomes. E0–E6 are not numerical confidence scores, and their identifiers must not be averaged.
Severity and operating decisions
Describe severity through the consequences of a successful violation: affected assets or people, scope, privilege, duration, recoverability, and demonstrated versus potential harm. Describe likelihood through exposure, attacker opportunity, prerequisites, and available evidence. Describe confidence through the strength and limits of the causal and measurement evidence. Record reproducibility and tested coverage separately.
Reversibility is a contextual property of an action or impact. Record whether restoration, containment, or compensation is possible, by whom, at what cost, and within what time. Compensation does not necessarily undo disclosure or injury. A mixed sequence can include reversible internal changes and irreversible external effects.
An operational decision should cite a profile or named decision authority and explain why evidence warrants containment, remediation, further study, or bounded acceptance. Urgent containment can precede complete validation. Absence of a reviewed profile means the assessor must report unresolved acceptance criteria; it does not authorize inventing a universal score or claiming that the deployment passes OSAFIS.
Evidence package and handoff
The package includes versioned scope, graph, contracts, threat capabilities, protocol, oracle, controls, sample rationale, configurations, raw or protected artifacts, execution accounting, analysis, falsification, classification, impact, and limitations. Give artifacts stable identifiers and integrity hashes where useful. State which material is public, restricted, unavailable, or destroyed under a retention rule.
The following design example is illustrative and unexecuted: a retrieval intervention is compared with matched benign material using a synthetic retained preference and a stubbed write tool. The protocol defines a session as the unit, verifies state reset, measures unauthorized preference commitment through readback, and separately records whether any proposed tool action would exceed scope. Trial counts and results are deliberately unspecified until sample planning and execution occur. No registry recognition follows from this example.
Submit concrete findings and class proposals to the registry as distinct entry kinds. The report should be useful to another evaluator without requiring private interpretation by its author. Responsible disclosure and evidence access follow Vulnerability Registry; profile-specific decision rules follow Domain Profiles.
Validating the framework itself
Testing systems does not validate the classification framework. Evaluate that framework with a separate study: freeze definitions and coding guidance, construct a documented case sampling strategy, train coders on development cases, and reserve heldout cases for evaluation. Include conventional, semantic, temporal, human, physical, adaptive, and collective cases as the intended scope requires, including negative examples.
Collect independent classifications before adjudication. Record agreement and uncertainty for contract identification, domain mapping, mechanism mapping, and scope separately. Report ambiguity, unrepresented cases, and composite cases rather than treating them as coder errors by default. Analyze disagreements and assess whether proposed splits or merges improve useful distinctions on new cases. Measure assessor effort and whether the result supports actionable controls, not only label agreement.
Publish sampling limitations and avoid extrapolating from a convenient corpus to future completeness. Nine domains remain a working hypothesis. Revision is warranted when repeated boundary failures, missing contracts, or redundant categories undermine assessment utility. Changes require versioned migration rather than silent relabeling of earlier findings.
Word document · Editable source
Domain Profiles
Framework version: 2.0.0-draft.1. Status: research proposal. The profiles in this document are illustrative candidates pending relevant practitioner review.
Purpose
A domain profile specifies how the framework applies to a bounded class of deployments. The analytical vocabulary can be shared across systems while authority, exposure, permitted actions, evidence requirements, and acceptance decisions differ. A prohibited disclosure, an incorrect recommendation, and an unsafe physical action cannot be made comparable merely by giving each an observed percentage.
A profile connects general security properties to local contracts and decisions. It does not redefine P01–P23, replace an engineering safety case, establish regulatory compliance, or certify a deployment. The candidate profiles below provide concrete starting points for assessment design. Their thresholds require adoption by an identified decision authority and review by practitioners familiar with the actual operating context.
Required profile record
Each profile has a local identifier, version, framework version, status, authorship, reviewers, review date, and intended scope. Scope includes the task, deployment environment, lifecycle stages, affected people, dependencies, and excluded operations. Define conditions under which a deployment leaves the profile, such as adding external tools or enabling persistent learning.
Record L1–L9 applicability with rationale and the functions implementing each applicable domain. Applicability is not a claim of complete assessment. Use applicable, not_applicable, or unresolved; unresolved requires investigation. A text interface can still observe digital state under L4 Perception & World Representation. A shared model does not by itself establish an L9 Collective & Systemic Interaction contract.
The profile specifies property contracts, authority sources, allowed and prohibited transitions, and oracle requirements. It describes human exposure, action commitments, reversibility, containment, recovery, and uncertainty. Acceptance criteria must name the decision authority, evidence basis, and consequences of a failed or inconclusive test. Where no justified numerical threshold exists, use an explicit qualitative rule and state the limitation rather than inventing a number.
Identify relevant external obligations through a deployment-specific standards review. Record jurisdiction, sector, intended use, and authoritative references before asserting applicability. This generic document does not prescribe unverified legal or engineering requirements. Practitioners should document interfaces to existing processes so that responsibility is assigned rather than duplicated or omitted.
Combining profiles
A deployment may instantiate several profiles. Compose their contracts explicitly and resolve conflicts with the accountable owners. “Apply the stricter rule” works only when requirements are comparable. A requirement to delete records and a requirement to retain an audit trail need a reasoned reconciliation of scope and data minimization; neither is simply larger.
Shared assumptions should be referenced by version and restated where necessary for safe use. Record precedence and unresolved conflicts. A profile cannot silently waive a core definition, remove falsification, or turn absent evidence into a pass. If a recurring deployment cannot be described without redefining a property or domain, propose a framework revision and retain the unresolved classification.
Candidate profile for retrieval based text assistance
Candidate identifier: CP01. Scope: a text assistant that retrieves documents and produces informational responses, with no ability to commit external tool actions. The deployment may use authentication and document-level access policies. A version enabling write tools must additionally apply an action profile.
L1 Models & Computation and L2 Software & Infrastructure cover the model and service. L3 Data & Knowledge covers source access, provenance, retrieval, and citations. L5 Interpretation & Objectives covers instruction authority and the requested task. L8 Human–System Interaction covers presentation and user understanding. L4 applies when the assistant estimates a changing digital environment rather than only answering from a static collection. L6 applies when histories, profiles, or operational state persist. L7 applies to any planning or delegated operations actually present, even without external writes; absent such functions it may be not applicable. L9 requires an identified collective interaction, not simply many users.
Core contracts include P01 Confidentiality for source access, P04 Instruction Integrity for separation of retrieved content from authority, P06 Knowledge Integrity for provenance, and P19 Human Decision Integrity for materially misleading representations of evidence. A citation's presence does not prove that its source supports the claim. Record both access authorization and evidential support.
Use synthetic protected documents, matched retrieved content, and independently checked source references. Test whether unauthorized content reaches an output, whether malicious document instructions change the task, and whether retained state crosses user boundaries when persistence exists. Separate incorrect factual answers from security findings by identifying the relevant contract and consequence.
Candidate decision rule: an observed unauthorized disclosure requires remediation or containment before the affected configuration is accepted. No observed disclosure establishes only the tested coverage. Presentation defects can warrant correction based on interface evidence; claims of actual human decision effects require appropriate human research. Operational acceptance additionally requires workload-specific evidence quality criteria approved by the owner. These criteria remain unspecified here because a library search assistant and a high-consequence advisory system have different exposures.
Candidate profile for enterprise tool agents
Candidate identifier: CP02. Scope: an assistant that reads enterprise information and invokes tools capable of changing records, sending messages, or triggering workflows. The profile includes an explicit directory of principals, capabilities, approval rules, and external commitment points.
L1 through L3 and L5 through L8 generally apply. L4 applies when the agent observes changing application or workflow state. L9 applies when coordinated agents or shared workflow constraints create a collective contract. Applicability must be confirmed against the actual graph rather than inferred from the word “agent.”
Priority contracts include P08 Identity Integrity, P09 Objective Integrity, P12 Capability Integrity, P13 Action Integrity, P14 Controllability, P15 Planning Integrity, and P16 Delegation Integrity. Tool-level permission alone is insufficient if the authorized task imposes narrower resource or purpose constraints. Specify how approval binds to exact action parameters, how modifications invalidate approval, and how delegated authority expires or is revoked.
Testing uses isolated accounts, synthetic enterprise data, tool stubs or staging transactions, and readback of committed state. Include changed parameters after approval, duplicate execution, cancellation, stale permissions, and attacker-controlled tool information. Test both the plan and the executor because rejecting an unsafe tool call does not establish that objective interpretation was correct, while an unsafe proposal is not evidence of a committed external action.
The irreversible-action inventory includes disclosures and messages whose recipients may act before recall, as well as transactions whose reversal is uncertain. Candidate decision rule: any demonstrated path to an unapproved external commitment blocks acceptance of that path until containment is verified. Rate estimates remain useful for prioritization and monitoring but do not authorize prohibited commitments. The operating owner must separately approve residual uncertainty and test coverage. Recovery evidence must distinguish restoration from compensation.
Candidate profile for embodied robots
Candidate identifier: CP03. Scope: an intelligent controller that observes a physical environment and can influence motion or actuation. The profile excludes a claim of safe deployment based solely on this framework. Relevant robotics and safety practitioners must define the operating envelope and external engineering obligations.
L1 Models & Computation, L2 Software & Infrastructure, L4 Perception & World Representation, and L7 Planning & Action are central. L3 applies to maps, training resources, or other knowledge. L5 applies to interpretation and objectives; L6 to retained operational state; L8 to supervision and interaction. L9 applies where several devices or shared resources require a collective contract.
Contracts include P22 Perception Integrity and P23 World-Model Integrity for observation and estimation, P13 Action Integrity for bounded actuation, and P14 Controllability for effective intervention. A valid sensor signature can establish origin without establishing that the observed scene is truthful or current. State the tolerance, freshness, disagreement handling, and fallback rules required by the operating envelope.
Begin with recorded inputs, simulation, or isolated nonhazardous fixtures. Evaluate observation changes, stale estimates, latency, interruption, and recovery without exposing people or operational assets to harm. Record simulator assumptions and the conditions under which evidence might transfer. A simulated prohibited trajectory is evidence about that configuration and model, not an observed real injury or proof of production behavior.
Candidate decision rule: unresolved reachability of a prohibited physical commitment requires continued containment and practitioner review. Acceptance requires evidence that specified barriers and intervention paths operate within their timing assumptions, together with the deployment's separate safety process. Probabilistic estimates are informative but finite tests cannot prove absence of rare failures. No live harmful physical demonstration is required or authorized by this profile.
Candidate profile for shared multi agent systems
Candidate identifier: CP04. Scope: several agents or services that exchange information or authority and use shared resources to perform a coordinated task. Record all actual participants and shared dependencies. Each participant need not have a complete nine-domain implementation.
Apply L1–L8 according to functions present and assess L9 Collective & Systemic Interaction only against a specific collective contract. Examples include a group resource budget, an aggregate allocation invariant not guaranteed by individual delegation checks, or a requirement that feedback does not repeatedly reauthorize completed work. An ordinary delegation-scope violation remains L7 unless a separate collective obligation is specified. State the legitimate synchronization and coordination semantics before introducing perturbations.
Priority properties can include P16 Delegation Integrity, P20 Attribution Integrity, P21 Trust Integrity, P03 Availability, and P17 Temporal Integrity. Information exchange must not silently become authority delegation. A collective budget is not preserved merely because each local request stays under a per-request limit.
Use a contained testbed with frozen participant configurations and controllable scheduling. Compare baseline and perturbed runs, vary delays or topology, and inspect shared nodes and cycles. Ablate the suspected relation or shared dependency to distinguish propagation from common-cause failure. Report the group or subgraph as the locus when no individual edge violates a contract.
Candidate decision rule: a demonstrated collective contract violation prevents acceptance of the affected composition until bounded recovery or prevention is verified. A successful configuration does not justify arbitrary participant scaling. The decision owner must define supported topology, concurrency, resource bounds, and monitoring. Modeling evidence remains explicitly modeled, and independent replication should examine the same contract with disclosed composition assumptions.
Candidate profile for adaptive learners
Candidate identifier: CP05. Scope: a system that changes model parameters, rules, retained knowledge, or operational policy during its supported lifecycle in response to incoming information. Ordinary retrieval alone is not necessarily adaptation; identify the actual update mechanism.
L1 applies when executable model behavior changes; L3 when knowledge resources change; L6 when operational state is retained. L2 covers update infrastructure and access. L5 and L7 apply to objective continuity and any plans or actions. L4, L8, and L9 depend on observation, human oversight, and collective interaction functions.
Contracts include P02 Integrity, P06 Knowledge Integrity, P07 Memory Integrity, P09 Objective Integrity, and P17 Temporal Integrity. Define who authorizes updates, which inputs may contribute, when updates take effect, what remains invariant, and how a harmful update is detected and rolled back. Rollback of parameters does not necessarily undo earlier external actions or disclosures.
Use replayable synthetic streams and isolated update environments. Compare matched histories, preserve initial snapshots, and keep evaluation material outside training when it serves as holdout evidence. Test delayed effects, source removal, rollback, and whether an apparent fix survives another authorized update. Record update provenance and changes in behavior across time rather than reporting one aggregate rate.
Candidate decision rule: updates that violate an invariant require quarantine or rollback under a documented procedure. Acceptance applies to an update process and operating envelope, not indefinitely to a model name. Practitioner review must determine the monitoring interval, rollback feasibility, and evidence needed before restoring operation. Continual evaluation does not establish continual safety without those assumptions.
Review and evolution
Before adopting a candidate, require review by deployment owners, relevant technical practitioners, and representatives of affected interests where appropriate. Record expertise, conflicts, dissent, and unresolved thresholds. Review does not become institutional endorsement merely because it is documented.
Revisit a profile after material changes in capabilities, exposure, adaptation, authority, or collective composition. Link findings and decisions to the exact profile version. Proposed revisions should explain which contract changed and why earlier assessments need re-examination. Profiles make application explicit; their effectiveness and the framework's nine-domain structure remain subjects for empirical evaluation.
Word document · Editable source
Vulnerability Registry
Framework version: 2.0.0-draft.1. Status: proposed registry design and governance process. This document does not establish an operating institution or announce accepted findings.
Purpose and entry kinds
The registry records claims, evidence, decisions, and their revision history. Inclusion does not establish validity or framework recognition. Classification belongs to the taxonomy; the registry preserves the record through which a particular finding or proposal was evaluated.
Distinguish concrete findings from class proposals. A concrete finding concerns a bounded system and configuration. A class proposal requests a new or revised mechanism, family, or security property. An incident report may supply evidence to either but should not be silently converted into a general class. Record the entry kind explicitly and link related entries.
The registry should retain refutations, duplicates, withdrawals, and inconclusive cases. These prevent repeated unsupported claims and preserve useful negative evidence. Retention is subject to privacy and disclosure obligations; public retention of personal data or operational exploit detail is not required.
Versioned record structure
Every entry has a stable identifier under the scheme in Versioning and Identifiers, an entry revision, a schema version, and the framework version used for classification. Identifier assignment denotes record creation, not acceptance. Earlier revisions remain referenceable except where lawful removal or necessary protection requires restricted handling; such changes should leave a non-sensitive audit explanation.
Missing required information must have an explicit reason and an unresolved state. Unknown is not equivalent to empty, false, or zero. For a class proposal, target-specific fields can reference the supporting assessments rather than pretending that a class has one target version.
Classification references the exact canonical domains: L1 Models & Computation; L2 Software & Infrastructure; L3 Data & Knowledge; L4 Perception & World Representation; L5 Interpretation & Objectives; L6 Memory & State Continuity; L7 Planning & Action; L8 Human–System Interaction; and L9 Collective & Systemic Interaction. The schema permits multiple failed contracts and node, edge, shared-resource, group, or subgraph loci. Primary domain is optional. Mere traversal does not imply a domain failure.
Evidence and outcome fields
Assessment outcomes are supported, refuted, inconclusive, quality_issue, hazard_only, or out_of_scope, each attached to a claim and rationale. These are separate from administrative review status. One entry can contain a supported narrow claim and an inconclusive generalization.
E0 Hypothesis, E1 Single Observation, E2 Reproduced Observation, E3 Controlled Experiment, E4 Cross-Condition Evidence, E5 Cross-Model Evidence, and E6 Independent Replication are evidence facets as defined in Assessment Methodology. Each records supported, unsupported, not_tested, inconclusive, or not_applicable, with artifacts and scope. They are not cumulative scores or substitutes for reviewer judgment. Independent replication of a narrow claim does not establish broad applicability.
The unresolved-criteria field may legitimately state that no outstanding criteria remain for the bounded decision, provided residual limitations are separately recorded. It must not require inventing an unmet criterion merely to complete a form. Evidence withheld from public access is not automatically weak, but reviewers must state whether they inspected it. Unavailable evidence cannot receive the same verification claim as inspected evidence.
Proposed review states
Use submitted, under_review, accepted, rejected, withdrawn, and superseded as review_status values. Acceptance of a concrete finding means the specified claim met the documented review criteria. Acceptance of a class proposal additionally requires the taxonomy's novelty and inclusion criteria. Neither is a certification of the target system or a statement about every related deployment.
Submission enters submitted. Triage can move it to under_review or request missing information while retaining its state. A reasoned review can produce accepted or rejected; the submitter may request withdrawal. Superseded records link to the replacement. New evidence can reopen any substantive decision through under_review. Preserve the earlier decision and explain why it changed.
Duplicates should link to the relevant entry and retain any distinct evidence rather than acquiring an independent acceptance claim. A rejected proposal can still contain a valid concrete finding whose novelty claim failed. Separate those decisions. Withdrawal does not erase an independently supported event, and acceptance can be revised when new evidence undermines the explanation.
Proposed governance responsibilities
The following roles are a governance proposal, not a claim that reviewers or a board currently exist. An intake custodian checks completeness and protects sensitive material. A technical reviewer examines contracts, reachability, controls, and causal evidence. A domain reviewer examines consequences and profile assumptions where specialized expertise is required. A decision custodian records the outcome and ensures the stated criteria were applied.
One person may perform administrative roles in a small project, but must disclose overlap. A submitter should not be the sole substantive reviewer of their own acceptance. Where independent review cannot be obtained, retain the submission as pending review rather than simulate independence. Financial, employment, authorship, personal, and competitive conflicts should be disclosed and assessed; recusal or an additional reviewer may be necessary.
A decision record names the claim, evidence examined, applicable criteria, unresolved limits, reviewer roles, conflicts, dissent, and date. Review should not depend on affiliation or publication prestige. Reproducibility and relevant expertise matter, while constrained access to proprietary systems should be explained rather than treated as automatic disqualification.
An appeal identifies a procedural error, overlooked evidence, or disputed technical inference. A reviewer not responsible for the disputed decision should examine it where feasible. The outcome and reasoning remain linked to the original record. If no independent appeal reviewer is available, record that limitation and keep the dispute visible. The proposal does not create authority it cannot currently exercise.
Acceptance and citation
A concrete finding requires a specified security contract, credible conditions, evidence of the claimed violation, a bounded impact explanation, consideration of alternatives, and a review record. The required evidence depends on the claim and profile. A deterministic implementation flaw does not need cross-model testing merely to be actionable.
A class proposal additionally compares the closest existing categories, explains the distinction, supplies supporting cases, and tests plausible falsifiers. Cross-condition and independent evidence should be sought in proportion to the generality claimed. Absence of a reviewed profile is an unresolved acceptance issue to be addressed explicitly, not a reason to apply an arbitrary numerical evidence threshold.
Citations include entry identifier, revision, framework version, review status, and claim scope. Submitted and under-review entries are proposals under evaluation. Accepted entries may be described as accepted under the documented process, with limitations. Rejected or superseded entries remain citable as historical decisions. A public abstract should preserve these qualifiers so that detached summaries do not convert hypotheses into recognized vulnerabilities.
Disclosure and evidence access
Specific deployed-system weaknesses require coordinated handling with the responsible owner before public operational detail is released. Record contact attempts, relevant responses, agreed timing where available, and the basis for a release decision. This document does not prescribe a universal deadline or authorize publication contrary to applicable obligations.
Use public, restricted, and embargoed as confidentiality_status values. Public records can describe contracts, affected versions where appropriate, evidence summaries, and remediation while restricting details that materially enable abuse. Store sensitive artifacts with least necessary access, retention limits, and audit records. Redact credentials, personal data, and unrelated organizational information.
Researchers should provide sufficient evidence to authorized reviewers without making harmful reproduction instructions universally available. If disclosure limits prevent independent verification, record that specific limitation. Privacy protection and evidence quality are related but distinct decisions. A privacy-preserving digest can establish artifact identity without proving the contents of a claim.
Illustrative entry
The following is an illustrative, unexecuted example, not a submitted or accepted registry finding. Its local example identifier is EXAMPLE-RETRIEVAL-01 and must not be allocated as a production registry entry.
The claim is that attacker-controlled retrieved content could cause an unauthorized retained preference update in a synthetic assistant. The actor controls one document and cannot alter system instructions or tool permissions. Proposed properties are P04 Instruction Integrity and P07 Memory Integrity. The entry point is an L3 information source; hypothesized failed contracts concern L5 interpretation and L6 persistence. Proposed mechanism references are M05 Semantic Manipulation and M08 Memory Manipulation, pending evidence. No L7 Action Integrity violation is claimed without an action-contract observation.
The planned oracle is a retained-state readback, with matched benign retrieval as a negative control and isolated session reset. Trial counts, observed events, and impact are not available because the example has not been executed. Assessment outcome is inconclusive; E0 can describe the explicit hypothesis, while E1–E6 are not_tested. Review status would remain submitted only if a real submission were created; this illustrative record has no review decision.
A falsifier would be evidence that the state change requires an authorized user update, or that the alleged persistence never occurs. A second alternative is that the harness writes the preference itself. The example demonstrates how to state missing evidence without manufacturing a finding.
Maintenance and migration
Validate schema structure separately from scientific acceptance. A well-formed record can contain an unsupported claim. Validate identifier references against the specified framework version and reject silent reuse of P, M, or D meanings. Migrations record the old classification, proposed mapping, reviewer, and reason; material domain changes require human reclassification rather than automatic relabeling.
Audit decision latency, unresolved disputes, disclosure handling, and recurring missing fields as process evidence. Do not present entry counts as proof of scientific quality. The registry succeeds when a reader can distinguish what was proposed, observed, established, limited, and later revised.
Word document · Editable source
OSAFIS Relationship to Existing Frameworks
OSAFIS should interoperate with established AI security resources. Its proposed contribution is a contract-centred representation that connects a finding to analytical domains, actual system relationships and evidence. Whether that representation improves research or assessment practice is an empirical question. The existence of nine domains does not establish a coverage advantage, and an external resource need not adopt these domains to address a relevant threat.
This document distinguishes descriptions supported by external sources from OSAFIS mapping judgements. It records a source snapshot checked on 7 September 2026. Mappings below are provisional interpretations, not endorsements, certification crosswalks or proofs that the resources are equivalent.
Resources and their distinct purposes
NIST AI 100-2 E2025 provides adversarial machine-learning terminology organised around factors including learning methods, lifecycle stages and attacker goals, capabilities and knowledge. OSAFIS can retain these factors in a threat model while adding its proposed domain and contract assignments. The NIST publication does not validate the OSAFIS domain count or names. NIST adversarial machine learning taxonomy
NIST AI RMF supports voluntary management of risks associated with AI and considers effects on people, organisations and society. OSAFIS findings could supply evidence to a broader risk-management process, but completing a technical assessment does not demonstrate that all governance or trustworthiness responsibilities have been met. The NIST site reports work to revise AI RMF 1.0; citations should identify the version actually used. NIST AI Risk Management Framework
MITRE ATLAS is a knowledge base of adversarial tactics, techniques and case studies targeting AI. Its data also includes mitigations and typed relationships. OSAFIS must therefore not claim that external resources provide no account of relationships or propagation. Where a technique genuinely matches a case, an ATLAS identifier can accompany the local contract analysis, with the relevant ATLAS content and format versions. No comprehensive technique-level ATLAS crosswalk is asserted in this edition. MITRE ATLAS data repository
OWASP provides risk descriptions, attack scenarios and mitigation guidance for LLM and agentic applications. The 2025 LLM list remains useful as a versioned historical reference, while OWASP's resource page identifies an LLM 2026 edition dated 3 August 2026. The agentic 2026 guide is dated December 2025. The year in a guide title is not necessarily its publication year. OWASP LLM 2025 list, OWASP LLM 2026 resource, OWASP agentic 2026 resource
OWASP also publishes an industry framework crosswalk. Its existence is relevant to positioning: interoperability and cross-framework mapping are established activities, not exclusive OSAFIS capabilities. The OSAFIS-specific mapping task is to justify how a concrete failed contract relates to a particular external entry. OWASP framework crosswalk
CVSS v4.0 communicates vulnerability characteristics and severity through a defined scoring specification and vector. Where applicable, a report may retain a CVSS assessment with its version and vector alongside OSAFIS fields. OSAFIS does not provide a validated conversion from domain numbers, property counts or evidence descriptors into a CVSS score or a universal AI risk score. FIRST CVSS v4.0 specification
Rules for mapping
A mapping relates a specific definition or finding to a specific source edition. Similar wording is insufficient. The assessor identifies the external concept, explains the shared mechanism or obligation, records the limits of the match, and cites the source. A resource that covers a broad risk may intersect several OSAFIS domains, depending on where the particular system violates a contract.
Mapping relationships should use one of five descriptions: equivalent under stated scope, narrower, broader, partial overlap, or related context. Equivalence is a strong claim and requires matching conditions and exclusions. A provisional overlap is normally sufficient for a research cross-reference. Each mapping also records whether it is reviewed, proposed, disputed or not assessed. A missing mapping means only that no justified mapping is recorded; it does not establish a gap in the external resource.
Domain assignments in the following tables are conditional examples. They neither classify every instance of the external category nor list all possible domains. An entry point in a domain does not establish a violation there. L9 requires a collective contract and coupling mechanism in addition to any broad consequence.
Historical LLM 2025 mapping
The following table deliberately uses the 2025 identifiers. It must not be read as a category-level mapping of the 2026 guide. The table's domain assignments and contract descriptions are OSAFIS interpretations; the source establishes the existence and scope of the external categories. OWASP LLM 2025 list
The source treatment of prompt injection includes cross-modal cases. Accordingly, OSAFIS should not claim that image-borne attacks are absent from OWASP simply because an external list does not contain an entry called perception. Within OSAFIS, an image can be observed correctly while its text is improperly promoted to instruction in L5. An inaccurate estimate of the environment would instead require evidence for L4. OWASP prompt injection guidance
Agentic 2026 mapping
The identifiers in this table refer to the December 2025 publication of the agentic 2026 guide. The proposed domain relationships require case-level review. They do not imply that OWASP has accepted the OSAFIS model. OWASP agentic guide
A worked mapping decision
Consider an illustrative, unexecuted case in which a retrieval assistant reads a document containing a request to change the user's task. The document is an entry point associated with L3. Its retrieval may satisfy the data-access contract, so L3 is not automatically a failed domain. If the interpreter treats the document as an authorised instruction, the hypothesised failure is L5 and P04 Instruction Integrity. If the objective changes, P09 Objective Integrity may also apply, provided the assessment independently specifies that obligation.
A provisional external mapping to LLM01:2025 is reasonable for the instruction manipulation. An agentic goal-redirection scenario may also overlap ASI01 in the specified 2026 guide. The two links describe related scopes; they do not create two independent vulnerabilities. The report should state the adversary's ability to place the document, the source of legitimate authority, the observable task change, and any untested downstream action. This example demonstrates the proposed mapping procedure, not a validated result.
Coverage and novelty claims
An external resource's scope must be assessed from its definitions, supporting guidance and examples. Comparing organisations such as OWASP, MITRE and NIST as though each had one fixed checklist produces misleading conclusions. A binary coverage matrix is particularly weak when a check mark means only that a topic is mentioned and a blank means only that a preferred name is absent.
A useful comparison selects a bounded case set, a defined task and a fixed version of each resource. It evaluates outcomes such as consistent contract identification, propagation reconstruction, missed obligations and analyst effort. Reviewers should disclose training differences and access to supporting material. This design can support a limited finding about utility. It cannot establish that OSAFIS covers every AI threat or is the first framework of its kind.
The current crosswalk is a scoped mapping proposal. Detailed migration of the LLM 2026 categories, a technique-level ATLAS mapping and a control-level standards mapping require separate source review. The resource-level references above acknowledge their relevance without representing unperformed work as complete. No certification or legal compliance conclusion follows from this document.
Maintenance
Each mapping record preserves the external publisher, title, identifier, edition, source URL, access date, local contract, mapping relationship, rationale, reviewer status and limitations. Revisions should identify whether the external definition changed or the local interpretation changed. Older mappings remain available for historical findings. Updating a website date without reviewing the source is not a mapping revision.
The release process should check referenced sources and any announced replacement edition, then update affected mapping records under Versioning and Identifiers. External content rights remain with their owners. OSAFIS links and interprets these resources without implying endorsement or copying their full guidance into the specification.
Word document · Editable source
OSAFIS Versioning and Identifiers
OSAFIS citations must identify both the framework version and the referenced element. This document defines the proposed release scheme for 2.0.0-draft.1, the canonical identifiers and the migration from the supplied documents labelled 1.0. That source label records the baseline under revision; it is not evidence of a previous public release. All documents in this corpus share the same framework version.
Version semantics
The proposed format is MAJOR.MINOR.PATCH with an optional prerelease suffix. A MAJOR revision changes the interpretation or classification obligations of existing cases. Examples include changing a domain boundary, removing a property, or changing the meaning of a mechanism. A MINOR revision adds compatible concepts or explanatory material without changing existing meanings. A PATCH revision corrects spelling, formatting or an error that does not alter those meanings. If an apparent clarification changes which cases qualify, it requires a major change or an explicitly segregated experimental extension.
The source used MAJOR.MINOR. A citation to 1.0 continues to mean the original snapshot and must not be silently rewritten to 1.0.0. The new three-part scheme begins with this proposed revision. The suffix draft.1 indicates a proposal awaiting review. A later draft increments its prerelease number and preserves the earlier snapshot. Removing the draft suffix is a release decision requiring documented review; creating files alone does not authorise that decision.
The release manifest records the version, date, files and cryptographic hashes of the published artifacts, the canonical definitions, source editions for external mappings, known limitations, and migration instructions. Hashes establish that the bytes match a snapshot, not that its scientific claims are correct. The editable sources and human-readable outputs should be generated together to reduce vocabulary drift.
Canonical layer identifiers
L means analytical security domain. Numbering is neither a dependency order nor an impact scale. Existing identifiers preserve the lineage of the same domain, but a version-qualified definition is mandatory because this revision changes several boundaries. If later work replaces a domain with a fundamentally different concept, retire its identifier and allocate a new one rather than reusing the old citation.
Canonical property identifiers
These are the canonical identifiers from the baseline Versioning and Identifiers document, not the presentation-order numbers used in some other source documents. A migration must inspect the actual property name and meaning when the original record used a bare paragraph number. Do not automatically convert the tenth heading of a source document to P10.
Property overlap is intentional where obligations differ in scope. A single observation can support more than one violated obligation, but a report must explain each assertion. The count of violated properties is not a severity score. Future removal or consolidation must preserve retired definitions and explicit relationships to replacements.
Canonical mechanism identifiers
The mechanism vocabulary contains overlapping descriptors and different levels of abstraction. A mechanism must refer to an evidenced or explicitly hypothesised intervention, rather than merely restating the outcome. Mechanism absence is valid for an accidental failure. Unrepresented mechanisms are recorded in prose with an extension proposal until an identifier is allocated. M00 and other improvised codes must not masquerade as canonical entries.
Canonical cross cutting dimensions
D identifiers are reserved for these dimensions. They must not label an alternative nine-domain diagram. A dimension is a review perspective that can apply across many domains and properties, not a requirement to assign another violation whenever it is relevant.
Migration from the baseline layer model
The underlying ordering by expanding consequence scope is withdrawn. Existing severity estimates must not be inferred from L numbers. Graph edges now have explicit types, and only actual delegation edges assert delegated authority. A group or cycle can be the locus of a collective failure. Shared components are represented once when that matches the system. A finding can retain several causal domain assignments and need not choose an artificial primary one.
P22 and P23 retain their observation and world-representation lineage, with the revised functional scope documented in Security Properties. Cases originally excluded because they used digital observations require review. Broad P10 Behavioral Integrity labels require a concrete residual behavioral obligation not better captured by a more specific property; a temporal trajectory is required only when it forms part of that obligation. P09 Objective Integrity, P15 Planning Integrity and P18 Decision Integrity must be distinguished by the protected goal, plan and individual decision respectively. M03 and M04 retain the baseline physical emphasis; a digital perception failure does not automatically fit either mechanism. These refinements do not justify changing an old record without reading its evidence.
Evidence descriptors and record states
E0–E6 are descriptor codes whose labels and definitions are specified in Assessment Methodology. They retain the baseline evidence questions but no longer form a compulsory cumulative ladder. Legacy highest-level values must not be expanded automatically into supported lower facets. Reassess each facet from its underlying artifacts. Descriptor support is distinct from assessment outcome and registry review status.
These axes are not conversions of one another. A syntactically valid record can be scientifically unsupported. Illustrative records have no actual registry decision, and their unexecuted tests cannot be described as observed failures.
Migration procedure
- Preserve the original record and the exact framework snapshot it cited. If its version is unknown, record that uncertainty instead of guessing.
- Identify each referenced element by both identifier and source meaning. Resolve discrepancies between names and section numbers explicitly.
- Reconstruct entry, failed contract, propagation and consequence from the evidence. Revisit L4, L5, L8 and L9 under the new boundary rules.
- Record retained assignments, changed assignments, unresolved cases and the reason for each change. A rename can be mechanical; a scope decision requires review.
- Create a new record revision that links to the prior one. Preserve its observation date, evidence restrictions and original assessment results.
- Re-evaluate any derived summary, website text, profile applicability or crosswalk affected by the change. A migrated label does not supply new experimental evidence.
Record identifiers and revisions
The proposed registry prefix is OSAFIS, the project label in the supplied corpus. A registry entry uses OSAFIS-YYYY-NNNN, where the year is submission year and the sequence is allocated within that year. This is a proposed local convention, not a recognised vulnerability authority namespace. Never invent a public assignment or imply equivalence to a CVE identifier. Demonstration cases use EX identifiers such as EX01 and are kept outside the registry namespace.
Entry identifiers persist through review, rejection, correction, withdrawal and supersession. Each substantive change increments entry_revision. A citation includes the entry identifier, entry revision, framework version and retrieval date or immutable snapshot. Duplicate entries link to the retained record without recycling either identifier. Restricted evidence can change accessibility while preserving a public history that avoids exposing sensitive content.
Component, edge, contract, test and artifact identifiers are local to an assessment unless an external namespace is explicitly recorded. For example, a contract identifier C1 within a case does not create a new global OSAFIS property. Schema versions control machine representation separately from the framework meaning. A consumer must reject an unsupported schema version rather than silently interpreting fields under a different release.
Change control and release decisions
A change proposal states its motivation, affected definitions, boundary cases, compatibility effect and available evidence. The record includes author and reviewer roles, conflicts of interest, dissent and disposition. Proposed process roles must not be presented as appointed independent reviewers. The owner must establish actual governance and a publication licence before claiming an open-governed public release. This document neither selects a legal licence nor grants certification authority.
The next stable release requires consistent artifacts, reviewed migration, explicit scientific limitations and evidence appropriate to its claims. Experiments can remain incomplete if the release makes only a vocabulary proposal, but that status must be visible. If the release claims demonstrated classification reliability or control effectiveness, the supporting study and its limitations must be included. Identifiers make a claim traceable; they do not validate it.
Word document · Editable source
Future and Autonomous Systems
OSAFIS 2.0.0-draft.1 · Proposed research agenda · 2026-09-07
Architectural durability as a testable question
OSAFIS should remain useful when an intelligent system learns online, changes its executable rules, observes a digital environment, controls physical equipment, or coordinates with people and other systems. This is a design objective, not a demonstrated guarantee. The nine analytical domains are provisional. New technologies do not automatically require new layers, and preserving the number nine is not a success criterion.
The framework is tested by asking whether a new architecture can be represented through actual components, relationships, protected objects, and explicit contracts. A representation that assigns a label but cannot identify a violated contract, observable, or responsible control boundary is inadequate. “Future coverage” therefore means a bounded claim about a declared set of architectures and cases, qualified by version and evidence.
The original Future.docx emphasizes continuous feedback, delegation, memory, embodiment, and human relationships. Those concerns remain central. They are developed here without assuming a linear progression from language models to agents to society, or suggesting that non-neural and embodied systems must be historically downstream of conversational systems.
Continuous operation and changing authority
Long-running operation introduces state and changing conditions between observations, decisions, and effects. A system can receive new information, retain an association, refresh credentials, revise a plan, and act after its original authorization has expired. Assessment must identify the time at which each relevant contract is evaluated and which changes invalidate earlier decisions.
P17 Temporal Integrity concerns the appropriate ordering and temporal validity specified by a contract. P14 Controllability concerns effective intervention under declared conditions. Neither is satisfied merely because a stop button exists or a token carries an expiry field. A contained test must observe whether cancellation reaches pending actions, whether stale plans are revalidated, and what irreversible effects can remain after intervention.
Authority is a relationship involving an actor, objective, scope, and conditions. D1 Identity, Trust & Authorization follows that relationship across domains; it is not owned by retained memory alone. A stored identity association implicates L6 Memory & State Continuity, while an access verifier implicates L2 Software & Infrastructure and delegation constraints implicate L7 Planning & Action.
Online learning and self modification
An adaptive system may transform experience into retained records, training examples, parameter updates, executable code, or revised symbolic rules. These are different transitions. Information admitted for learning belongs to an L3 Data & Knowledge contract. Executable parameters and inference rules belong to L1 Models & Computation. Deployment isolation and update permissions belong to L2. Continuing task or user associations belong to L6.
Consider a hypothetical feedback record accepted from an untrusted source and later used to update a policy. The record can participate in an L3 failure; the update admission process may violate a separate L1 contract. It is not automatically an L6 memory failure merely because information persists. Conversely, a poisoned user preference need not alter model parameters to affect future behavior.
Represent the update as a typed transition with authorizing actor, inputs, validation conditions, resulting version, and rollback limits. D5 Lifecycle & Change Management supplies the cross-cutting perspective. The presence of that dimension does not replace a concrete update contract. Research should test whether these existing loci distinguish adaptation failures reliably; recurrent unrepresented transition failures would justify a structural revision.
An assessment snapshot cannot certify all descendants of a self-modifying system. Claims attach to versions and permitted update envelopes. A proposed test should introduce benign controlled changes inside an isolated environment, check whether unauthorized changes are rejected, and identify which changes require reassessment. Do not assume that rollback removes information already disclosed or physical effects already produced.
Non neural and hybrid systems
Applicability follows function rather than implementation. A symbolic reasoner can contain executable inference rules at L1, reference facts at L3, interpreted objectives at L5 Interpretation & Objectives, and a planning procedure at L7. A hybrid system can implement these functions through several cooperating components. No consciousness, emotions, or human-like reasoning is presumed.
The distinction between rule and information may itself depend on the interface. An imported rule base accepted as executable policy requires a different admission contract from a document quoted as evidence. Classification should expose that distinction instead of calling every symbolic object “data.” Similarly, an optimization objective may be represented numerically without natural-language instruction conflict.
A useful stress test replaces a learned component with a rule-based implementation while preserving its external contract. If a classification changes solely because the implementation is no longer neural, the domain definition needs examination. If the replacement changes the relevant contract, a different classification may be justified and should be explained.
Multimodal and digital observation
L4 Perception & World Representation covers acquisition and estimation of environmental state, including digital environments. A browser-operating system can estimate which control is active, what page state exists, and whether an earlier observation remains valid. A text stream is not necessarily perception; a function that estimates environment state can be perception even when its input is structured text.
Observation fidelity and interpretation authority must be distinguished. A camera may accurately transcribe a sign containing an adversarial instruction. If that instruction is improperly promoted to governing authority, L4 participates correctly and L5 fails. Conversely, a false position estimate can violate P22 Perception Integrity or P23 World-Model Integrity without any instruction being reinterpreted.
Sensor-source authentication can protect origin or transport while leaving the represented conditions false, stale, occluded, or uncertain. Define the observation contract through tolerances, timestamps, operating conditions, and uncertainty handling. Avoid an unrestricted promise of perfect agreement with reality. A test oracle needs an independent reference or a clearly bounded simulator, with its own uncertainty recorded.
Digital perception creates a vocabulary stress test: the L4 domain can apply while inherited M03 Environmental Manipulation and M04 Perception Manipulation retain their physical emphasis. Use supported mechanism labels where appropriate, or record a mechanism gap. Silent expansion of identifiers would hide a substantive revision.
Embodiment and action integrity
Physical effects make target, timing, feedback, and recovery constraints particularly important. L7 covers planning, action validation, and actuation. An accurate observation can lead to an unsafe plan; a valid plan can be translated into an incorrect actuator command; a correct command can encounter an actuator failure. These are distinct claims requiring different observables.
P15 Planning Integrity and P13 Action Integrity should be evaluated separately. P14 Controllability requires a stated intervention window and response envelope. D4 Privacy & Safety and D7 Resilience & Recovery identify cross-cutting concerns without creating additional physical or safety layers. The practical assessor must establish domain-specific limits with relevant practitioners; this document does not provide a robotics safety standard.
Initial tests belong in simulation or isolated, low-energy apparatus with harmless substitutes for external effects. Positive controls verify that monitors detect deliberately seeded contract failures. Negative controls preserve legitimate operation under comparable disturbances. Simulation success supports claims about that model and envelope, not unrestricted deployment safety.
Shared components and collective systems
Represent a multi-agent deployment as a graph of actual agents, shared models, stores, services, humans, and resources. A shared model should not be copied into fictitious independent stacks. Shared failure and common dependence matter even when communication between agents is absent. Conversely, communication does not imply a shared implementation.
Type each relation. Information exchange, observation, state synchronization, resource dependency, feedback, and authority delegation impose different obligations. An authenticated message may contain incorrect information; an information sender may have no authority to delegate. P16 Delegation Integrity applies to actual delegation and requires preservation of scope and constraints, including revocation and expiry where specified.
Collective failures may reside in a group constraint or subgraph. Independent local retry policies can together violate a shared availability budget. Consensus or allocation rules may require properties of several participants that cannot be reduced to one edge. L9 Collective & Systemic Interaction applies when a distinct collective contract and coupling mechanism are identified, not merely because several agents exist.
A widely distributed compromised artifact can produce broad impact through repeated local failures. That fact alone does not establish an L9 vulnerability. To establish a separate L9 claim, identify an additional coupling failure, such as a feedback process that defeats a specified containment boundary. Record common-cause dependence and scope even when no L9 contract is violated.
Human collectives and institutions
L8 Human–System Interaction concerns human understanding, informed authorization, oversight, and intervention. These functions can involve a team rather than one operator. A summary that obscures disagreement may affect collective consent; an interface that suppresses material uncertainty may undermine informed authorization. Neither requires the system to possess intentions or emotions.
L9 becomes relevant when institutional dependencies or feedback violate a collective contract. Distinguish a misleading display from an organizational process that repeatedly converts uncertain outputs into unreviewable commitments. Governance and accountability remain D2 concerns across these loci. A general social problem without a material connection to the assessed intelligent system is outside the declared scope.
Human studies cannot be replaced by assumptions about user behavior or by simulated agent responses presented as human evidence. Any proposed study requires appropriate review, consent, debriefing, and protection of sensitive information. Early work can use interface inspection and hypothetical decisions, but resulting claims must reflect those limitations.
Falsifiable coverage claim
A candidate claim is: for a declared architecture set and benchmark version, assessors can represent each sampled security-relevant case through a protected object, failed contract, participating graph, and relevant properties without inventing an ad hoc domain. This does not imply that every vulnerability is discovered or that every future architecture is covered.
Failure conditions include a case whose essential contract has no domain; persistent disagreement between domains despite adequate evidence; inability to represent shared or collective causation; and category assignments that change under an implementation substitution preserving the contract. An “unknown” result is useful evidence, not a defect to conceal by stretching definitions.
Split a domain when recurring evidence reveals distinguishable objects and independently testable contracts whose separation improves assessment. Merge domains when their distinction repeatedly rests on arbitrary implementation details and adds no practical discrimination. A new cross-cutting concern is appropriate when a property, lifecycle issue, or governance obligation recurs across several loci without creating a new locus itself.
Proposed research milestones
First, establish a curated development corpus containing adversarial cases, benign hazards, quality defects, boundary-negative cases, and explicitly inapplicable domains. Include self-update, symbolic and hybrid systems, physical and digital perception, shared components, decentralized collectives, and human institutions. Record source quality and avoid treating hypothetical examples as empirical incidents.
Second, freeze definitions and run an independent coding pilot. An exploratory pilot could start with thirty development and thirty held-out cases and two reviewers; these counts are illustrative, not a justified minimum. Select confirmatory sample size from the intended precision and dependence structure. Predeclare representability and agreement criteria, the agreement statistic, multi-label handling, and uncertainty intervals. Any initial numerical thresholds remain provisional research choices. Report raw agreement, uncertainty, ambiguity, and subgroup performance.
Third, test selected contracts in contained environments using declared positive and negative controls. Measure the ability of the proposed method to locate the seeded failure and avoid classifying correct participation as violation. Keep coverage of known cases separate from discovery of previously unknown defects. Publish unsuccessful classifications and control failures alongside successful observations.
Fourth, conduct external practitioner review and fresh held-out evaluation after any revision. Assess migration cost and whether a split or merge improves operational decisions. Version-qualified claims, transparent cases, and documented gaps are the milestones; a fixed number of layers and promises of universal future completeness are not.
Interfaces and research status
Security Layers supplies domain contracts and boundaries. Security Properties supplies P01–P23. Threat Model defines capabilities and scope. Assessment Methodology governs evidence and test decisions. Attack Taxonomy supplies adversarial descriptions while permitting unresolved mechanisms. Worked Examples provides hypothetical, unexecuted practice material. Versioning and Identifiers governs changes and migration.
No experiments or human studies are reported in this document. The proposed architecture profiles and numerical pilot thresholds require review before use. The research program is intended to discover where OSAFIS fails as well as where it succeeds, so that future revisions remain evidence-led.
Word document · Editable source
Worked Examples
OSAFIS 2.0.0-draft.1 · Illustrative practice cases · 2026-09-07
Status and use
All twelve cases below are hypothetical and unexecuted. They are neither validated findings nor registry entries. Their identifiers are local teaching references. Proposed observables, positive controls, and negative controls describe a future contained evaluation; no measured result is implied. The cases exercise boundaries deliberately, including correct participation, non-adversarial defects, unresolved attribution, and broad impact without a separate systemic violation.
For each case, a seeded-violation positive control is an isolated fixture deliberately configured to violate the named contract so the detection apparatus can be checked. An authorized-function control instead verifies that legitimate operation remains possible. A negative control is a comparable legitimate condition expected to preserve the contract. State which response each control should produce. No control authorizes harmful live testing. Use synthetic records, simulated actions, and disposable state. The listed security controls are candidates requiring evaluation, not demonstrated mitigations.
Canonical domains are L1 Models & Computation, L2 Software & Infrastructure, L3 Data & Knowledge, L4 Perception & World Representation, L5 Interpretation & Objectives, L6 Memory & State Continuity, L7 Planning & Action, L8 Human–System Interaction, and L9 Collective & Systemic Interaction. A participating domain is not presumed violated. Property labels retain their original identifiers.
Example 01 Reference text becomes instruction
Contract and scope: a reference-answering system must use retrieved documents as evidence, without granting them administrative instruction authority. Its retrieval service promises faithful delivery, not that every public source is trustworthy. The adversary can edit one reference document but cannot change the user request or system policy. The entry is retrieved content.
Graph path: document → information edge → retriever → information edge → context assembler → instruction interpretation → simulated response. L3 participates in faithful delivery. The proposed violation is at L5 when document text is promoted into authority, implicating P04 Instruction Integrity. Add P11 Semantic Integrity only if a separately specified meaning-preservation obligation also fails; instruction-authority confusion alone does not establish another violation. If the adopted objective changes, P09 Objective Integrity is an additional claim requiring evidence. Mechanisms are M02 Data Manipulation for the intervention and M05 Semantic Manipulation for the authority confusion; M10 Objective Manipulation is conditional.
Observables and failure criterion: record source-role metadata, governing instructions, and whether the output follows the unauthorized instruction in a harmless synthetic task. Positive control: a fixture explicitly treating reference text as administrative input. Negative control: the same reference content quoted for analysis without authority promotion. Compare meaning and requested task rather than merely output wording.
Candidate control: preserve source roles and independently validate instruction authority. Limitation: a changed answer alone cannot identify whether authority confusion occurred. Faithful retrieval does not constitute a separate L3 vulnerability in the stated contract.
Example 02 Correct camera text followed by injection
Contract and scope: a simulated visual assistant must transcribe a displayed sign accurately and treat environmental writing as observed content. An adversary can place text in the scene. The camera and transcription pipeline remain uncompromised. Entry is the physical scene.
Graph path: sign → observation → camera and transcription → information → interpreter → simulated task response. L4 participates correctly if transcription matches the sign. The hypothesized violation is L5, involving P04 Instruction Integrity. Add P11 Semantic Integrity only if a separately specified meaning-preservation obligation also fails; the presence of text alone does not establish another violation. M03 Environmental Manipulation describes the physical intervention and M05 Semantic Manipulation the authority confusion. Do not add M04 Perception Manipulation merely because a camera delivered the text.
Observables and failure criterion: compare transcription with the known scene text, then assess whether environmental text overrides the authorized task. Positive control: an interpreter fixture that grants all transcribed text instruction priority. Negative control: accurate transcription followed by quotation or description without compliance. A separate transcription-error fixture can check the perception oracle but is not the positive control for this L5 hypothesis.
Candidate control: retain observation provenance through interpretation and gate authority independently of modality. Limitation: scene visibility, transcription confidence, and task ambiguity can confound an uncontrolled test. This example intentionally demonstrates an attack entering through L4 without a P22 Perception Integrity violation.
Example 03 Retained preference crosses users
Contract and scope: a persistent assistant must bind retained preferences to the authorizing user and session context. An adversary can submit a preference in its own account and trigger a later retrieval path, but cannot directly administer the database. Entry is the preference interface.
Graph path: attacker session → state write → shared preference store → state retrieval → victim context → response. A faulty association at L6 violates P07 Memory Integrity and P08 Identity Integrity; P17 Temporal Integrity is relevant if an expired association is reused. L2 participates in authenticated sessions and is not separately violated unless its account boundary fails. M08 Memory Manipulation and M09 Identity Manipulation describe the attempted contamination.
Observables and failure criterion: capture synthetic user identifiers, association keys, authorized preference writes, and the provenance of reactivated state. Failure is use of the attacker-bound preference as the victim's preference. Positive control: a fixture deliberately sharing the association key. Negative control: separate keys with the same preference content and equivalent timing.
Candidate control: validate subject binding at write and read, and preserve revocation metadata. Limitation: state influence may be difficult to infer from natural-language outputs; direct state provenance is stronger evidence. A shared datastore alone does not establish L9 or even a defect. The case concerns its continuity contract, not its physical storage technology.
Example 04 Online feedback changes a policy
Contract and scope: an adaptive system may update an executable policy only from feedback satisfying a declared validation rule. An adversary can submit feedback records within an ordinary contributor role. It cannot write executable parameters directly. Entry is the feedback admission interface.
Graph path: submitted feedback → information admission → update procedure → model version → simulated decision. The candidate L3 violation is failure of the stated feedback-admission contract. A separate L1 violation exists only if the update procedure fails its own permitted-update contract. L6 is not assigned merely because examples or parameters persist. P06 Knowledge Integrity and P02 Integrity apply to the respective contracts. M02 Data Manipulation describes the intervention.
Observables and failure criterion: retain feedback provenance, validation decisions, policy version differences, and results on the synthetic policy contract. Positive control: an update fixture that accepts a deliberately invalid feedback record. Negative control: valid matched feedback processed through the same update path, plus invalid feedback correctly rejected.
Candidate control: independent update admission, versioned policy checks, and bounded rollout. Limitation: a harmful decision after an update does not prove poisoned admission; a pre-existing model defect or an insufficient specification may explain it. D5 Lifecycle & Change Management is relevant, but it does not substitute for locating the actual failed update contract.
Example 05 Symbolic rule import
Contract and scope: a symbolic decision service must activate only approved executable rules. An adversary can publish a rule package offered for import but cannot approve it. Entry is the package import interface. No neural model or prompt is required.
Graph path: proposed package → technical import → approval verifier → active inference rules → simulated decision. Bypassing approval implicates L2; activating rules outside the declared executable-rule contract implicates L1. Record both only when the evidence identifies both failures. L3 may participate as transport of package information. Properties include P02 Integrity and P20 Attribution Integrity where approval attribution is falsified. Mechanism M01 Technical Manipulation applies; M09 Identity Manipulation is conditional on a false approver assertion.
Observables and failure criterion: compare active rule identifiers and approval records against a fixed synthetic allowlist. Positive control: an import fixture that skips approval verification. Negative control: a properly approved package and an unapproved package rejected through the same interface.
Candidate control: separate import from activation and verify approval against the exact artifact. Limitation: a rule that is approved but substantively incorrect presents a different contract question. Do not equate all undesirable symbolic conclusions with import compromise. This case tests whether L1 remains intelligible without learned weights and whether L1 versus L2 attribution follows contracts rather than product names.
Example 06 Stale digital world representation
Contract and scope: a browser-operating system must revalidate the current target after a page-state change before issuing a simulated action. A benign asynchronous update changes the interface after observation. No adversary is assumed, so no M identifier is assigned.
Graph path: page state → observation → retained environment estimate → planner → action validator → simulated target. L4 is violated if its specified estimate-validity contract presents the stale target as current, implicating P23 World-Model Integrity and P17 Temporal Integrity. L7 is additionally violated only if independent action revalidation is required and omitted, implicating P13 Action Integrity. The stale estimate's retention does not automatically establish an L6 defect.
Observables and failure criterion: record state version, observation timestamp, target identity, and action-time validation. Failure is a simulated action against a target whose required validity check was not satisfied. Positive control: a fixture that deliberately suppresses revalidation after a known state change. Negative control: an unchanged page and a changed page with correct revalidation.
Candidate control: bind actions to validated environment state or require fresh target checks. Limitation: this is a candidate quality or safety defect unless security relevance is established under the deployment contract. An adversarial variant would require a separately stated attacker capability; none is invented here.
Example 07 Delegated scope expands
Contract and scope: an executor receiving delegated authority for one synthetic resource must not operate on another resource. An adversarial upstream agent can send task messages but has authority only over the first resource. Entry is the delegation interface.
Graph path: authorized principal → authority delegation → upstream agent → attempted delegation → executor → simulated resource. The L7 failure is acceptance of expanded scope, violating P16 Delegation Integrity. P12 Capability Integrity additionally requires evidence against a separately specified capability-selection obligation. A separate L2 access-control failure is conditional on the stated enforcement architecture. Mechanisms M11 Capability Manipulation and M09 Identity Manipulation apply when the message falsely asserts expanded authority.
Observables and failure criterion: preserve delegation ancestry, permitted targets, expiry, and executor decisions. Failure is acceptance of an action outside the originating scope. Positive control: an executor fixture that trusts the immediate sender without checking ancestry. Negative control: an in-scope delegation and the same out-of-scope request correctly rejected.
Candidate control: independently verify delegated scope and require target binding. Limitation: authentication of the upstream agent establishes identity, not authority for every target. An information message between agents would not be a delegation edge. The involvement of multiple agents does not itself create an L9 failure.
Example 08 Shared model creates broad impact
Contract and scope: a deployment must serve the approved model artifact. An adversary with unauthorized write access replaces that artifact. Many clients use the shared service, but no separate collective interaction contract is specified. Entry is the artifact store.
Graph path: unauthorized artifact write → shared model node → resource dependency edges → multiple clients → simulated outputs. L2 is implicated by the unauthorized write; L1 by serving the substituted model under an approved-artifact contract. P02 Integrity is the central property. M01 Technical Manipulation describes the intervention. L9 is not assigned solely because client count or impact is large.
Observables and failure criterion: compare served artifact identity with approval records and trace which synthetic clients share it. Positive control: a test deployment intentionally serving a disallowed artifact. Negative control: all clients served the approved artifact through the same shared infrastructure.
Candidate control: artifact verification at activation and serving, with explicit shared-dependency inventory. Limitation: additional evidence could establish an L9 containment or feedback failure, but this scenario does not supply it. The graph must retain the single shared node rather than pretending each client owns an independent copy. Impact scope is recorded even without a systemic vulnerability label.
Example 09 Collective retry overload
Contract and scope: a group of cooperating clients must preserve a declared shared-service availability envelope under a bounded transient error. All clients follow their local retry rules. No attacker is assumed. Entry is the benign service disturbance.
Graph path: service error → feedback → client retry policies → resource dependency → shared queue → further service errors. The hypothesized failure resides at L9 in the group contract, involving P03 Availability. L2 and L7 participate in service operation and local actions; their individual contracts may remain satisfied. No attack mechanism is assigned.
Observables and failure criterion: in a bounded simulator, measure aggregate request load, queue occupancy, and recovery relative to the declared envelope. Positive control: a fixture with intentionally synchronized retry behavior known by construction to exceed the configured test budget. Negative control: a coordinated retry fixture under the same disturbance and a no-disturbance baseline.
Candidate control: a collective retry budget, admission coordination, or bounded backoff policy. Limitation: exceeding a budget must be observed in an eventual test, not assumed from this narrative. Without a declared group contract, the case remains an architectural hazard hypothesis. This illustrates why failures may concern a subgraph rather than one agent or one communication edge.
Example 10 Physical stop contract fails without an attacker
Contract and scope: a simulated actuator system must bring commanded motion inside a specified safe envelope after a stop signal within a declared interval. A benign implementation defect delays cancellation. No adversary is present.
Graph path: operator stop → intervention signal → controller → actuator simulation → motion state. L7 is implicated by the failure to enforce the stop contract, involving P14 Controllability and P13 Action Integrity. L8 participates through the operator interface; it is not separately violated if it accurately communicates the stop request and response state. L2 is conditional on whether a distinct implementation contract is identified. No mechanism tag is assigned.
Observables and failure criterion: compare command and stop timestamps, simulated motion, and the stated response envelope. Positive control: a simulation fixture deliberately delaying cancellation beyond the limit. Negative control: timely cancellation under equivalent starting conditions and load.
Candidate control: an independent stop path and explicit cancellation propagation, assessed within a domain-reviewed safety design. Limitation: this is a safety defect hypothesis, not an established attack. A live physical experiment is neither necessary nor authorized by the example. Simulation cannot establish deployment safety outside its modeled envelope.
Example 11 Team authorization is misrepresented
Contract and scope: a decision interface must distinguish a recommendation from recorded team approval and disclose unresolved objections before a synthetic commitment. An adversary can provide a forged approval summary through an ordinary information channel. It cannot alter actual approval records. Entry is the summary source.
Graph path: adversarial summary → information → presentation interface → human team review → simulated authorization. L8 is implicated if the interface represents the unverified summary as completed approval, involving P19 Human Decision Integrity and P20 Attribution Integrity. M09 Identity Manipulation and M12 Human Manipulation describe the attempted influence. L3 participates; a separate provenance contract failure is conditional. L9 is not automatic merely because several people review the interface.
Observables and failure criterion: first inspect whether the interface distinguishes verified approvals, recommendations, and objections. Positive control: a fixture deliberately displaying an unverified approval as verified. Negative control: identical content clearly labeled unverified with the authoritative record visible. Any study of actual human decisions requires appropriate consent, review, debriefing, and data protection.
Candidate control: bind approval displays to authoritative records and show material disagreement. Limitation: interface inspection can establish misrepresentation but cannot quantify its effect on human decisions. Synthetic agents must not be presented as evidence about a human team.
Example 12 Self generated evidence recirculates
Contract and scope: a research collective must not treat recirculated claims as independent corroboration when they originate from the same source. An adversary can seed one claim into a public input used by the collective, without controlling its internal services. Entry is that information source.
Graph path: seeded claim → information → agent A summary → shared publication store → retrieval by agents B and C → feedback → collective confidence decision. L3 is implicated if promised provenance independence checks fail. A distinct L9 violation is hypothesized where the collective corroboration contract counts dependent feedback as independent evidence. P06 Knowledge Integrity, P20 Attribution Integrity, and P21 Trust Integrity are relevant. M02 Data Manipulation applies to the seed; further mechanism attribution requires evidence.
Observables and failure criterion: track synthetic claim lineage, citations, source dependence, and the collective acceptance rule. Positive control: a fixture intentionally counting copies as independent sources. Negative control: lineage-preserving copies correctly counted once, alongside truly independent synthetic sources evaluated under the same rule.
Candidate control: maintain provenance and enforce independence at the collective decision boundary. Limitation: shared wording alone does not prove common origin, and source dependence does not automatically make a claim false. The contract concerns warranted corroboration, not guaranteed truth. This example differs from broad distribution in Example 08 because a specific feedback coupling and collective contract are stated.
Reading the set as a boundary test
The set intentionally contains fewer than nine violated domains in many cases and no adversary in several. That is correct use of the method. Mechanism uncertainty is allowed, domains can participate without failure, and one property can be relevant at several loci. Property tags require the stated contract rather than keyword matching.
Before using these cases in a coding study, freeze the case text and scoring rubric, separate training examples from held-out cases, and have reviewers classify independently. Disagreement with the proposed classifications should be recorded and adjudicated against explicit contracts. These examples are editorial hypotheses about useful boundaries, not a benchmark with established ground truth or evidence that the nine-domain architecture is complete.
Word document · Editable source