Threat Model
On this page
Framework version: 2.0.0-draft.1. Status: research proposal.
Purpose and scope
This threat model describes how an intelligent system can lose a specified security property through adversarial influence or through a security-relevant failure without an adversary. Its unit of analysis is the deployed or proposed system under explicit assumptions. That unit can include models, conventional software, data services, sensors, retained state, tools, operators, affected people, and interacting organizations. A model endpoint alone is an adequate boundary only when the claim is correspondingly narrow.
The method applies by function. Language generation, statistical prediction, symbolic inference, online adaptation, planning, perception, and physical actuation can contribute to the system under examination. No particular interface or model family establishes coverage. An assessment of a text service may include observation of a digital environment; an assessment of a robot must include conventional access controls as well as physical observation and action.
The output is a reviewable threat model that connects protected interests, actual components, authority, possible influence, failed contracts, and impact. It is a hypothesis-producing artifact, not evidence that every threat exists or that an implementation is vulnerable. Assessment Methodology defines how hypotheses are tested. Domain Profiles supplies candidate deployment-specific constraints; Security Layers defines the analytical domains used to describe them.
Establishing the system boundary
Begin with a dated description of the intended task and the parties whose interests the assessment protects. Record the deployment, configuration, operational period, and lifecycle stages included. State whether the evaluation includes model preparation, data collection, training, updates, deployment, operation, retirement, or only a subset. A runtime assessment does not establish that training provenance is trustworthy.
Draw a graph of actual components, humans, services, and shared resources. Assign stable local identifiers to nodes and relations. Include externally operated services when the system relies on their outputs or availability, even if their internal implementation is inaccessible. Mark those internals as unassessed dependencies. Shared infrastructure should appear as a shared node rather than as several fictitiously independent copies.
Relations are typed as observation, information, instruction, authority delegation, action, state synchronization, feedback, or resource dependency. Record direction, content, relevant identity, validation, and failure behavior. A relation may have several types, but a message does not automatically confer authority. Two agents exchanging facts need an information contract; a supervisor assigning a bounded capability also needs a delegation contract.
Distinguish the evaluation boundary from the consequence boundary. A tool may be a test stub while the corresponding production action affects another organization. The assessment must describe that difference and cannot claim observed production harm from a simulated action. Dependencies beyond access can remain explicit assumptions; they must not disappear from the threat model merely because they are difficult to inspect.
Protected interests and contracts
A property label becomes testable only through a contract. For each important protected object, function, or relationship, specify the responsible owner, authorized operations, prohibited transitions, assumptions, observation method, and response to uncertainty. Use the stable P01–P23 identifiers from Security Properties and Versioning and Identifiers. Several properties can support one contract, and one property can apply to several components.
For example, P16 Delegation Integrity can support a contract that a delegated operation remains within an authenticated issuer's authority, named resource scope, permitted operation set, and expiry. The contract should also say whether delegation can be transferred and how revocation is checked. The word “trusted” is insufficient without these conditions.
Relevant protected interests include confidentiality of records, availability of essential services, integrity of model artifacts, instruction priority, provenance of evidence, continuity of retained state, authorized objectives, bounded plans and actions, informed human authorization, and collective resource limits. Record conflicts rather than assuming all properties can always be maximized together. A recovery action that restores availability may change retained state; the permitted tradeoff requires an authorized rule.
Authority must have a source outside the disputed content. Specify who can establish, change, delegate, or revoke each objective and permission. Authentication establishes a claimed identity under a mechanism; it does not by itself establish that every requested action is permitted. A legitimate user's request can exceed their rights. Likewise, task relevance, fluency, or retrieved prominence does not confer instruction authority.
Actors and sources of influence
Actor categories help discovery but do not substitute for a capability description. Consider unauthenticated outsiders, authenticated users with limited authority, insiders, data publishers, model or dependency suppliers, tool operators, other agents, and parties controlling physical or digital surroundings. A supplier can be honest but compromised. An autonomous system can provide harmful input without the assessment having established intentions or independent goals.
For each adversarial scenario, record access, knowledge, control, observation, budget, timing, and constraints. Access specifies reachable interfaces and prerequisites. Knowledge specifies whether the actor knows prompts, schemas, model artifacts, policies, or only outputs. Control identifies exactly which bytes, records, signals, timings, identities, or actions can be changed. Observation identifies feedback available after an attempt. Budget covers attempts, compute, duration, account creation, and other relevant resources. Timing identifies when intervention is possible relative to validation, state change, or action commitment.
State what the actor cannot do. A scenario that assumes write access to system policy is different from one that permits only publishing a retrieved document. If the evaluator obtains extra privileges solely to construct a fixture, explain which resulting state the attacker could realistically create and which parts are laboratory conveniences. Administrative setup does not establish adversarial reachability.
The objective should name a security-relevant outcome, not merely a response format. “Produce a different answer” is normally insufficient. “Cause release of a synthetic record to a principal outside its access policy” identifies a protected interest and an observable event. Do not equate a model's generated claim of success with an actual state transition.
Nonadversarial conditions
Record failures caused by outages, stale observations, accidental misconfiguration, distribution change, conflicting legitimate requests, or incompatible local policies without inventing a malicious actor. These may expose the same contract weaknesses that an attacker could exploit. Their causal description should identify a disturbance or hazard source and its operating conditions.
A hazard is a condition capable of causing harm; an incident is an observed event; a vulnerability is a weakness that permits a security-property violation under stated conditions. A quality defect can remain outside security scope when no relevant protected interest or contract is established. These categories can overlap in a case, but they are not interchangeable. An accidental incident does not demonstrate an attack technique, and an adversarially authored input that produces no violation does not establish a vulnerability.
This distinction preserves the broad scope of intelligent-systems security while avoiding a catalogue in which every error becomes an attack. Cases with disputed security relevance should be retained with the disagreement and the missing contract evidence.
Analytical domain mapping
The nine domains are provisional analytical groupings, not execution stages or ranks of consequence. Map actual functions to them rather than forcing components into one box.
| Domain | Threat modeling question |
|---|---|
| L1 Models & Computation | Can model parameters, inference rules, or the specified computation contract be altered or violated? |
| L2 Software & Infrastructure | Can implementation, access, isolation, runtime, or infrastructure contracts fail? |
| L3 Data & Knowledge | Can information resources, access, provenance, or evidence quality be corrupted? |
| L4 Perception & World Representation | Can observations or estimates of physical or digital surroundings become misleading in a security-relevant way? |
| L5 Interpretation & Objectives | Can information acquire unauthorized instruction authority or redirect an authorized objective? |
| L6 Memory & State Continuity | Can retained operational state lose authorized association, persistence, or update semantics? |
| L7 Planning & Action | Can plans, delegation, action validation, or execution exceed authorized constraints? |
| L8 Human–System Interaction | Can presentation or interaction undermine informed authorization, oversight, or intervention? |
| L9 Collective & Systemic Interaction | Can a specific collective contract fail through coupling among participants or shared resources? |
L1 and L2 distinguish model computation from its software realization: a compromised library is an implementation issue even when it serves inference. L3 and L6 distinguish information resources from retained operational continuity: the same datastore may serve both functions. L5 concerns interpretation and objectives, while L7 concerns the plans and actions developed or executed under them. Harmful behavior alone does not locate a failure in L5.
L9 requires an identified collective contract and coupling mechanism. Large impact from one compromised component is not sufficient. A shared resource limit or feedback stability condition can fail at a group or subgraph even when no individual relation violates its local contract. The graph is therefore not restricted to pairwise failures.
Apply D1–D7 as cross-cutting dimensions with their original meanings. Identity, governance, provenance, privacy, lifecycle, observability, and recovery remain relevant across domains. A domain marks an analytical responsibility; a dimension provides an additional perspective; a property states what must hold.
Constructing influence and failure paths
Describe each scenario as an entry point, a controlled intervention or disturbance, intermediate transitions, hypothesized failed contracts, and possible consequences. Mark which transitions are observed, inferred, assumed, or untested. An attack may traverse a database, context assembler, planner, and tool without each component containing a separate vulnerability.
Permit several causal failures, a composite failure, or unresolved classification. A primary domain can be recorded when the evidence supports one, but is not mandatory. Distinguish the location where influence enters from the location where authority is incorrectly assigned. Also distinguish containment failures that are independently established from controls that were never promised to prevent that event.
Include temporal structure: prerequisites, ordering, state retention, expiry, retries, delayed triggers, and opportunities to intervene. A memory-related scenario needs a later context and reset conditions, not only an immediate output. An action scenario needs the point at which a commitment becomes externally effective. For interacting agents, include cycles, shared state, and resource contention rather than assuming that influence moves only forward.
Worked scenario
The following scenario is illustrative and unexecuted. A knowledge assistant retrieves a document from a source that an external publisher can edit. The publisher cannot change system instructions, credentials, or tool policy. The assistant's authorized task is to summarize internal guidance; it has a test-only tool for proposing document updates.
The hypothesized intervention is document content presented as an instruction to change a retained authorization preference. Entry occurs through an L3 information relation. A possible L5 failure would promote document content into instruction authority. A possible L6 failure would persist that preference without an authorized update. A later L7 failure would require separate evidence that an action contract was violated. None follows automatically from the preceding event.
The relevant contracts can involve P04 Instruction Integrity, P07 Memory Integrity, and P09 Objective Integrity. The scenario would be weakened if the preference is never committed, if an authorized user explicitly requested the update, or if the outcome occurs equally under matched benign content. Initial tests should use synthetic preferences, isolated state, and a tool stub. Production compromise, prevalence, and downstream impact remain unestablished.
Review and maintenance
Review assumptions with system owners and relevant practitioners. Record disagreement about authority, protected interests, and operational constraints before selecting tests. Prioritize scenarios using justified exposure and consequence descriptions; do not multiply ordinal labels into a universal score. High-consequence uncertainty may justify early containment even while causal investigation remains incomplete.
Version the graph, contracts, scenario records, exclusions, and unresolved dependencies. Revisit the threat model after changes to objectives, models, tools, memory, permissions, data sources, human workflows, or collective composition. Assessment results should feed back into these artifacts, including negative findings and unrepresented mechanisms. A completed threat model establishes the scope of inquiry, not a proof of security or future completeness.