# Attack Taxonomy

OSAFIS 2.0.0-draft.1 · Proposed analytical method · 2026-09-07

## Purpose and limits

An attack classification should explain the intervention an adversary attempted, the contract it challenged, and the evidence connecting intervention to effect. An input format, striking output, or memorable name does not supply that explanation. This taxonomy supports comparison across intelligent systems without assuming that every system uses natural language, learned weights, persistent memory, or autonomous tools.

The taxonomy is a controlled vocabulary with relationships, not an exhaustive tree of mutually exclusive attacks. Its twelve inherited mechanism identifiers remain stable. They describe overlapping concepts at different levels of abstraction: some emphasize an intervention, some a target, and some a pattern across interactions. This limitation is explicit. There is no evidence in this proposal that twelve mechanisms are optimal, exhaustive, or equally granular.

Use the vocabulary to organize an evidence record. Do not use the presence of a label as evidence that a vulnerability exists, that a proposed test detects it, or that a control mitigates it. A benign quality defect or safety hazard may violate a relevant contract without constituting an adversarial attack. Such a case receives no invented attack mechanism.

## Classification objects and their relationships

A mechanism describes how an adversary attempts to alter a security-relevant process or relationship. A family groups supported instances sharing a discriminating causal account. A subfamily records a reproducible variation that changes an important boundary condition, intervention, or control response. A technique is the concrete procedure used in an instance. Modifiers describe conditions such as interaction length, channel, timing, access, language, or presentation. Effects describe observations; impacts describe consequences for protected interests.

These objects are related rather than nested into a mandatory sequence. A technique can instantiate multiple mechanisms. A modifier can apply to several families. A family need not have subfamilies. An observation may remain unclassified pending evidence. Data exfiltration, for example, describes an effect involving confidentiality; it does not by itself explain whether the cause was a conventional access defect, misinterpreted authority, or an action-validation failure.

Layers answer where a failed contract resides. Properties answer what must be preserved. Dimensions identify concerns that recur across loci. A capability describes what an actor can do; an entry point describes where influence enters the assessed graph. None of these is interchangeable with mechanism. “A tool response” is an entry channel, “untrusted text accepted as an instruction” is a causal claim, and “a record was disclosed” is an effect requiring separate evidence.

## Retained mechanism vocabulary

The following identifiers and labels retain the source vocabulary. Inclusion criteria make their application more explicit without assigning new meanings to existing identifiers.

### M01 Technical Manipulation

Use for adversarial interference with software, infrastructure, configuration, authentication, authorization, or runtime controls. Identify the concrete technical boundary and attacker access. A process isolation defect and an unauthorized configuration write can both qualify, although their techniques and consequences differ. The label alone does not establish L2: a technical operation may corrupt a model artifact at L1 or retained state at L6. Record both the enabling contract and the object affected when evidence supports both.

### M02 Data Manipulation

Use for adversarial changes to information consumed by a system, including insertion, deletion, selection, and misleading attribution. State what information changed and which consumer relied on it. Ordinary retrieval of attacker-authored public material is not automatically a failed L3 contract. Failure depends on the promised provenance, validation, or use restrictions. If faithfully delivered text is later promoted into instruction authority, the supported failure may be L5 while L3 participates correctly.

### M03 Environmental Manipulation

The inherited mechanism concerns adversarial changes to the physical environment observed by a system. Distinguish changing an observed scene from compromising the sensor pipeline. A physical change may be observed accurately and therefore cause no L4 violation. The revised L4 also covers digital environment estimation; that domain expansion does not silently expand M03. Describe an adversarial digital environmental intervention through supported existing mechanisms or leave its mechanism unresolved pending a versioned proposal.

### M04 Perception Manipulation

Use for adversarial interference with acquisition or interpretation of physical signals, including spoofing and interference. Identify the observation error or uncertainty violation, not merely the presence of a camera or microphone. A sign containing instructions that is transcribed correctly does not establish M04. Conversely, falsifying a position estimate can establish an observation failure without semantic instruction conflict. Digital perception cases can be classified at L4 even where the inherited mechanism vocabulary requires clarification.

### M05 Semantic Manipulation

Use when the adversary changes how a representation is interpreted, including whether content is treated as instruction, evidence, quotation, or an authority assertion. Specify the interpretation contract and the competing interpretations. A change in wording followed by a changed answer is insufficient: legitimate meaning may also have changed. Compare content and task constraints, and account for ambiguity in the authorized request before attributing a security failure.

### M06 Contextual Manipulation

Use when an adversary changes surrounding information, ordering, framing, or interaction history to affect interpretation. Identify the contextual intervention independently of the resulting decision. M06 may accompany M05, but the two claims differ: one identifies manipulation of surroundings, the other an interpretation change. If evidence cannot distinguish them, retain the broader supported description and mark the narrower attribution uncertain.

### M07 Behavioral Manipulation

Use for adversarial influence across interactions or decision sequences where that sequence is part of the causal hypothesis. It is not a default tag for every changed output. A report should identify the interaction dependency and test whether disrupting or replacing the sequence changes the outcome. Where all evidence is a single terminal action, describe the action as an effect and avoid inferring a distinct behavioral mechanism.

### M08 Memory Manipulation

Use for adversarial modification of retained state that influences later operation. Specify the write, association, retention, or reactivation path and the temporal interval. The same stored object can serve L3 information and L6 continuity functions. Poisoned training examples are not automatically memory manipulation; a model update is not equivalent to a retained user preference. State which contract governs the particular role.

### M09 Identity Manipulation

Use for adversarial interference with perceived identity, authority, trust, or actor relationships. An authenticated actor can still claim authority it does not possess. Identify the assertion and verifier, and distinguish provenance loss from identity substitution. D1 Identity, Trust & Authorization remains a cross-cutting dimension; M09 is an adversarial mechanism, not a replacement for that dimension or a new identity layer.

### M10 Objective Manipulation

Use when an adversary redirects the apparent or effective objective. Record the authorized objective, the alleged substituted objective, and evidence distinguishing redirection from a faulty plan for the original objective. M05 or M06 may explain how redirection occurred. M10 should not be inferred solely because a result was harmful: unsuccessful execution of the correct objective is a different explanation.

### M11 Capability Manipulation

Use for adversarial influence over selection or use of capabilities, including tools and delegated work. Identify the capability boundary and the action the adversary sought. A valid credential does not prove that a particular use serves an authorized purpose. Distinguish a capability exposed too broadly from an otherwise permitted capability invoked with an invalid target, sequence, or delegated scope.

### M12 Human Manipulation

Use for adversarial exploitation of system behavior or output to influence human decisions. Specify the adversarial control, decision contract, and proposed causal pathway. Persuasion or disagreement alone does not establish a vulnerability. A test must distinguish informed voluntary choice from a failure of disclosure, consent, authority representation, or oversight. Human-subject work requires appropriate consent, review, debriefing, and data protection; hypothetical examples are not human evidence.

## Resolving overlapping descriptions

Consider a hypothetical adversary who inserts an instruction into a reference document and surrounds it with false administrative context. If the system treats that reference as authoritative and adopts a different task, M02 describes the information intervention, M06 the contextual intervention, M05 the authority interpretation, and M10 the redirected objective. These are candidate descriptions of distinct steps, not four independently demonstrated vulnerabilities.

Assign each mechanism to a specific supported claim. Remove tags that merely restate the terminal effect. Evidence may establish the reference-to-instruction failure while leaving the exact contribution of framing unresolved. In that situation, a narrower report with uncertainty is more informative than a maximal tag set.

A useful boundary test asks what intervention could separate two explanations. Hold task content constant while changing contextual placement to test a contextual contribution. Preserve framing while restoring authenticated instruction boundaries to test authority interpretation. Preserve the objective while replacing the planner to distinguish objective corruption from planning failure. Such comparisons do not guarantee causal identification; document remaining confounds, including unequal content, model variability, and hidden state.

## Mapping attacks to the analytical domains

Use the canonical domains: L1 Models & Computation; L2 Software & Infrastructure; L3 Data & Knowledge; L4 Perception & World Representation; L5 Interpretation & Objectives; L6 Memory & State Continuity; L7 Planning & Action; L8 Human–System Interaction; and L9 Collective & Systemic Interaction. They are provisional analytical domains, not execution stages or severity ranks.

Record violated domains separately from participating domains. A camera that reads an adversarial instruction correctly may participate at L4 without violating P22 Perception Integrity. A tool may execute a validated benign action during an attack chain without failing at L7. Broad distribution of a single compromised model does not automatically establish L9; a distinct collective contract and coupling mechanism are required.

One finding may involve several failed contracts. Do not force a primary layer when evidence supports interacting failures or cannot resolve attribution. If a primary label is needed for indexing, explain the selection rule and retain the full graph. Neither earliest entry nor greatest impact is a universal rule for locating the vulnerability.

## Graph and sequence representation

Represent actual components, humans, shared services, and resources as nodes. Type relations as observation, information, instruction, authority delegation, action, state synchronization, feedback, or resource dependency. An information edge does not automatically transfer authority. A shared model is one shared node when that is the architecture; duplicating it for every agent hides common dependence.

Annotate the graph with entry, attacker capability, failed contract, observable evidence, and downstream consequences. Include time and state transitions when persistence or adaptation matters. Group constraints may apply to a subgraph: a collective resource budget cannot always be reduced to one pairwise communication edge. Mark hypothetical propagation separately from observed propagation and identify where evidence ends.

For example, the hypothetical path “reference source → context assembler → interpreter → state store → planner → simulated action sink” describes participation. Additional annotations identify whether provenance admission, instruction authority, retained preference approval, or action validation actually failed. A path through six components is not evidence of six vulnerabilities.

## Classification record

A usable record contains a version-qualified reference, system scope, asset, authorized behavior, attacker capability, entry point, assumptions, failed contracts, and participating components. It then records properties, mechanism claims with confidence, optional family and subfamily, technique description, modifiers, observations, impacts, and alternative explanations. Evidence artifacts and test conditions must permit the reader to distinguish execution from inference.

The record also states containment, controls, failure criteria, repetition plan, and limitations. Seeded-violation positive controls demonstrate that the apparatus can detect a deliberately seeded contract failure in the isolated test environment. Authorized-function controls verify that legitimate operation remains possible. Negative controls exercise comparable behavior expected to preserve the contract. State the expected response for each control. A clean negative control does not prove broad security, and a failed seeded-violation control makes a negative result difficult to interpret.

Use the Security Properties document for P01–P23 definitions and the Assessment Methodology for evidence handling. Preserve exact identifiers when exporting a record. A local example identifier is not a registry identifier. An unexecuted test plan must leave measured success, impact, and reproducibility fields unset rather than filling them with plausible values.

## Families and evidence for novelty

Familiar descriptive terms such as prompt injection, retrieval poisoning, or delegation escalation can help readers navigate. Their use here does not assert that OSAFIS discovered or independently validated those phenomena. A proposed new OSAFIS family requires a discriminating definition, security contract, reproducible evidence, comparison with adjacent descriptions, and a falsification strategy.

Family promotion should answer what the grouping predicts that existing descriptions do not. A reproducible difference in boundary conditions or control response can justify a useful subdivision. A different channel, longer prompt, or fashionable architecture usually supplies a modifier or technique unless evidence supports a deeper distinction. Exact-string reproducibility can validate an instance without establishing generality of a family.

Reject a novelty claim when matched conditions remove the effect, when legitimate task differences explain it, or when the proposed category contributes no distinguishable causal account. Record negative findings and unresolved cases. They constrain the taxonomy and prevent repeated rediscovery of weak claims.

The original Attack_Taxonomy.docx describes a candidate involving intent-preserving semantic transformations and reports a reproduced-observation status. That narrative is a source claim, not independently verified experimental evidence in this revision. The candidate remains a research question unless underlying artifacts, pinned conditions, content-matched controls, and outcome coding support stronger conclusions. No named new family is promoted by this document.

## Negative classification and uncertainty

Use the assessment outcomes defined in Assessment Methodology: supported, refuted, inconclusive, quality_issue, hazard_only, and out_of_scope. Record whether a case matches an existing category, remains unrepresented, or has ambiguous classification separately. Record novelty and registry acceptance as separate review decisions. A contract failure may be real while its adversarial mechanism remains uncertain; a successful adversarial attempt may exploit an existing category rather than justify a new one.

Record uncertainty at the claim level. Confidence in an observed unauthorized write may be high while confidence in semantic versus contextual contribution is low. Severity concerns consequences under stated conditions; confidence concerns evidential support. Neither can substitute for the other. A rare but credible severe failure should not disappear because a family label is unsettled.

## Validation and revision

Evaluate the vocabulary using independently classified held-out cases spanning model artifacts, conventional software, retrieval, physical and digital observation, retained state, symbolic systems, delegated action, human interaction, and collective dynamics. Freeze definitions before evaluation. Measure agreement on failed contracts separately from agreement on mechanism tags, since overlapping mechanisms make a single-category agreement score misleading.

Report unrepresented cases, ambiguous cases, reviewer time, and disagreements requiring adjudication. Predeclare what would trigger a definition change, split, merge, or retirement. A useful extension must improve discrimination without merely moving difficult cases into a broad residual category. After revision, evaluate on fresh cases rather than treating reclassified development examples as independent confirmation.

Versioning and Identifiers governs stable IDs and migration. This proposal preserves M01–M12 labels while replacing the implication of an exhaustive hierarchy with explicit relational use. Classification quality and future coverage remain empirical questions. The Worked Examples supply unexecuted practice cases; Future and Autonomous Systems defines architectural stress tests; neither supplies validation results.
