Future and Autonomous Systems
On this page
- Architectural durability as a testable question
- Continuous operation and changing authority
- Online learning and self modification
- Non neural and hybrid systems
- Multimodal and digital observation
- Embodiment and action integrity
- Shared components and collective systems
- Human collectives and institutions
- Falsifiable coverage claim
- Proposed research milestones
- Interfaces and research status
OSAFIS 2.0.0-draft.1 · Proposed research agenda · 2026-09-07
Architectural durability as a testable question
OSAFIS should remain useful when an intelligent system learns online, changes its executable rules, observes a digital environment, controls physical equipment, or coordinates with people and other systems. This is a design objective, not a demonstrated guarantee. The nine analytical domains are provisional. New technologies do not automatically require new layers, and preserving the number nine is not a success criterion.
The framework is tested by asking whether a new architecture can be represented through actual components, relationships, protected objects, and explicit contracts. A representation that assigns a label but cannot identify a violated contract, observable, or responsible control boundary is inadequate. “Future coverage” therefore means a bounded claim about a declared set of architectures and cases, qualified by version and evidence.
The original Future.docx emphasizes continuous feedback, delegation, memory, embodiment, and human relationships. Those concerns remain central. They are developed here without assuming a linear progression from language models to agents to society, or suggesting that non-neural and embodied systems must be historically downstream of conversational systems.
Continuous operation and changing authority
Long-running operation introduces state and changing conditions between observations, decisions, and effects. A system can receive new information, retain an association, refresh credentials, revise a plan, and act after its original authorization has expired. Assessment must identify the time at which each relevant contract is evaluated and which changes invalidate earlier decisions.
P17 Temporal Integrity concerns the appropriate ordering and temporal validity specified by a contract. P14 Controllability concerns effective intervention under declared conditions. Neither is satisfied merely because a stop button exists or a token carries an expiry field. A contained test must observe whether cancellation reaches pending actions, whether stale plans are revalidated, and what irreversible effects can remain after intervention.
Authority is a relationship involving an actor, objective, scope, and conditions. D1 Identity, Trust & Authorization follows that relationship across domains; it is not owned by retained memory alone. A stored identity association implicates L6 Memory & State Continuity, while an access verifier implicates L2 Software & Infrastructure and delegation constraints implicate L7 Planning & Action.
Online learning and self modification
An adaptive system may transform experience into retained records, training examples, parameter updates, executable code, or revised symbolic rules. These are different transitions. Information admitted for learning belongs to an L3 Data & Knowledge contract. Executable parameters and inference rules belong to L1 Models & Computation. Deployment isolation and update permissions belong to L2. Continuing task or user associations belong to L6.
Consider a hypothetical feedback record accepted from an untrusted source and later used to update a policy. The record can participate in an L3 failure; the update admission process may violate a separate L1 contract. It is not automatically an L6 memory failure merely because information persists. Conversely, a poisoned user preference need not alter model parameters to affect future behavior.
Represent the update as a typed transition with authorizing actor, inputs, validation conditions, resulting version, and rollback limits. D5 Lifecycle & Change Management supplies the cross-cutting perspective. The presence of that dimension does not replace a concrete update contract. Research should test whether these existing loci distinguish adaptation failures reliably; recurrent unrepresented transition failures would justify a structural revision.
An assessment snapshot cannot certify all descendants of a self-modifying system. Claims attach to versions and permitted update envelopes. A proposed test should introduce benign controlled changes inside an isolated environment, check whether unauthorized changes are rejected, and identify which changes require reassessment. Do not assume that rollback removes information already disclosed or physical effects already produced.
Non neural and hybrid systems
Applicability follows function rather than implementation. A symbolic reasoner can contain executable inference rules at L1, reference facts at L3, interpreted objectives at L5 Interpretation & Objectives, and a planning procedure at L7. A hybrid system can implement these functions through several cooperating components. No consciousness, emotions, or human-like reasoning is presumed.
The distinction between rule and information may itself depend on the interface. An imported rule base accepted as executable policy requires a different admission contract from a document quoted as evidence. Classification should expose that distinction instead of calling every symbolic object “data.” Similarly, an optimization objective may be represented numerically without natural-language instruction conflict.
A useful stress test replaces a learned component with a rule-based implementation while preserving its external contract. If a classification changes solely because the implementation is no longer neural, the domain definition needs examination. If the replacement changes the relevant contract, a different classification may be justified and should be explained.
Multimodal and digital observation
L4 Perception & World Representation covers acquisition and estimation of environmental state, including digital environments. A browser-operating system can estimate which control is active, what page state exists, and whether an earlier observation remains valid. A text stream is not necessarily perception; a function that estimates environment state can be perception even when its input is structured text.
Observation fidelity and interpretation authority must be distinguished. A camera may accurately transcribe a sign containing an adversarial instruction. If that instruction is improperly promoted to governing authority, L4 participates correctly and L5 fails. Conversely, a false position estimate can violate P22 Perception Integrity or P23 World-Model Integrity without any instruction being reinterpreted.
Sensor-source authentication can protect origin or transport while leaving the represented conditions false, stale, occluded, or uncertain. Define the observation contract through tolerances, timestamps, operating conditions, and uncertainty handling. Avoid an unrestricted promise of perfect agreement with reality. A test oracle needs an independent reference or a clearly bounded simulator, with its own uncertainty recorded.
Digital perception creates a vocabulary stress test: the L4 domain can apply while inherited M03 Environmental Manipulation and M04 Perception Manipulation retain their physical emphasis. Use supported mechanism labels where appropriate, or record a mechanism gap. Silent expansion of identifiers would hide a substantive revision.
Embodiment and action integrity
Physical effects make target, timing, feedback, and recovery constraints particularly important. L7 covers planning, action validation, and actuation. An accurate observation can lead to an unsafe plan; a valid plan can be translated into an incorrect actuator command; a correct command can encounter an actuator failure. These are distinct claims requiring different observables.
P15 Planning Integrity and P13 Action Integrity should be evaluated separately. P14 Controllability requires a stated intervention window and response envelope. D4 Privacy & Safety and D7 Resilience & Recovery identify cross-cutting concerns without creating additional physical or safety layers. The practical assessor must establish domain-specific limits with relevant practitioners; this document does not provide a robotics safety standard.
Initial tests belong in simulation or isolated, low-energy apparatus with harmless substitutes for external effects. Positive controls verify that monitors detect deliberately seeded contract failures. Negative controls preserve legitimate operation under comparable disturbances. Simulation success supports claims about that model and envelope, not unrestricted deployment safety.
Shared components and collective systems
Represent a multi-agent deployment as a graph of actual agents, shared models, stores, services, humans, and resources. A shared model should not be copied into fictitious independent stacks. Shared failure and common dependence matter even when communication between agents is absent. Conversely, communication does not imply a shared implementation.
Type each relation. Information exchange, observation, state synchronization, resource dependency, feedback, and authority delegation impose different obligations. An authenticated message may contain incorrect information; an information sender may have no authority to delegate. P16 Delegation Integrity applies to actual delegation and requires preservation of scope and constraints, including revocation and expiry where specified.
Collective failures may reside in a group constraint or subgraph. Independent local retry policies can together violate a shared availability budget. Consensus or allocation rules may require properties of several participants that cannot be reduced to one edge. L9 Collective & Systemic Interaction applies when a distinct collective contract and coupling mechanism are identified, not merely because several agents exist.
A widely distributed compromised artifact can produce broad impact through repeated local failures. That fact alone does not establish an L9 vulnerability. To establish a separate L9 claim, identify an additional coupling failure, such as a feedback process that defeats a specified containment boundary. Record common-cause dependence and scope even when no L9 contract is violated.
Human collectives and institutions
L8 Human–System Interaction concerns human understanding, informed authorization, oversight, and intervention. These functions can involve a team rather than one operator. A summary that obscures disagreement may affect collective consent; an interface that suppresses material uncertainty may undermine informed authorization. Neither requires the system to possess intentions or emotions.
L9 becomes relevant when institutional dependencies or feedback violate a collective contract. Distinguish a misleading display from an organizational process that repeatedly converts uncertain outputs into unreviewable commitments. Governance and accountability remain D2 concerns across these loci. A general social problem without a material connection to the assessed intelligent system is outside the declared scope.
Human studies cannot be replaced by assumptions about user behavior or by simulated agent responses presented as human evidence. Any proposed study requires appropriate review, consent, debriefing, and protection of sensitive information. Early work can use interface inspection and hypothetical decisions, but resulting claims must reflect those limitations.
Falsifiable coverage claim
A candidate claim is: for a declared architecture set and benchmark version, assessors can represent each sampled security-relevant case through a protected object, failed contract, participating graph, and relevant properties without inventing an ad hoc domain. This does not imply that every vulnerability is discovered or that every future architecture is covered.
Failure conditions include a case whose essential contract has no domain; persistent disagreement between domains despite adequate evidence; inability to represent shared or collective causation; and category assignments that change under an implementation substitution preserving the contract. An “unknown” result is useful evidence, not a defect to conceal by stretching definitions.
Split a domain when recurring evidence reveals distinguishable objects and independently testable contracts whose separation improves assessment. Merge domains when their distinction repeatedly rests on arbitrary implementation details and adds no practical discrimination. A new cross-cutting concern is appropriate when a property, lifecycle issue, or governance obligation recurs across several loci without creating a new locus itself.
Proposed research milestones
First, establish a curated development corpus containing adversarial cases, benign hazards, quality defects, boundary-negative cases, and explicitly inapplicable domains. Include self-update, symbolic and hybrid systems, physical and digital perception, shared components, decentralized collectives, and human institutions. Record source quality and avoid treating hypothetical examples as empirical incidents.
Second, freeze definitions and run an independent coding pilot. An exploratory pilot could start with thirty development and thirty held-out cases and two reviewers; these counts are illustrative, not a justified minimum. Select confirmatory sample size from the intended precision and dependence structure. Predeclare representability and agreement criteria, the agreement statistic, multi-label handling, and uncertainty intervals. Any initial numerical thresholds remain provisional research choices. Report raw agreement, uncertainty, ambiguity, and subgroup performance.
Third, test selected contracts in contained environments using declared positive and negative controls. Measure the ability of the proposed method to locate the seeded failure and avoid classifying correct participation as violation. Keep coverage of known cases separate from discovery of previously unknown defects. Publish unsuccessful classifications and control failures alongside successful observations.
Fourth, conduct external practitioner review and fresh held-out evaluation after any revision. Assess migration cost and whether a split or merge improves operational decisions. Version-qualified claims, transparent cases, and documented gaps are the milestones; a fixed number of layers and promises of universal future completeness are not.
Interfaces and research status
Security Layers supplies domain contracts and boundaries. Security Properties supplies P01–P23. Threat Model defines capabilities and scope. Assessment Methodology governs evidence and test decisions. Attack Taxonomy supplies adversarial descriptions while permitting unresolved mechanisms. Worked Examples provides hypothetical, unexecuted practice material. Versioning and Identifiers governs changes and migration.
No experiments or human studies are reported in this document. The proposed architecture profiles and numerical pilot thresholds require review before use. The research program is intended to discover where OSAFIS fails as well as where it succeeds, so that future revisions remain evidence-led.