# CHI/CHIS Specification v0.1

Dmitry Chistyakov · 2026-09-20 · Research proposal and open technical specification

**Evidence status:** no human calibration, real model benchmark, or construct-validation study has been performed for this release. The reference code contains synthetic demonstrations only. CHI measures observable human-like cognitive dynamics, not subjective experience or moral status. It is not a consciousness, sentience, IQ, safety, or personhood measure.

## 1. Scope and conformance

This specification defines measurement objects, an initial 30-test panel, a run protocol and reporting rules. “MUST” denotes a reporting requirement for claims of conformance to v0.1. “SHOULD” denotes a recommended practice whose omission must be explained. A conforming **capacity-only** report may omit human similarity and CHI. A **human-referenced** report requires a frozen empirical human calibration and all five axes. A **synthetic demonstration** is a separate evidence class and may never be described as either empirical class.

The evaluated system includes the model, persistent stores, retrieval and update policy, tools, scheduler and coordination. Record all of them. Hidden chain-of-thought is neither required nor scored. The observable transition is `(S_next, output) = F(weights, policy, S, input, environment, randomness)`. S includes persistent state. Logged history is an external record, not a claim about private mental content.

The supplied Python code is a **scoring and logging starter kit**, not an implementation of all 30 task environments. For a real benchmark, researchers MUST additionally freeze and release exact item instances or generators, prompts, acceptable-answer keys, feature extractors, model adapters and calibration materials. Template counts below define a concrete pilot panel but do not establish power.

## 2. Two profiles and CHI

Required axes: R reasoning; M memory; C cognitive autonomy/endogenous cognitive dynamics; A affective modulation; I integration.

For every axis report: raw per-test outcomes; capacity p in [0,1]; a declared behavioral feature vector z; optional human similarity h in [0,1]; uncertainty; sample sizes; missingness; and measurement conditions. Capacity and human similarity MUST remain separate. Do not call normative accuracy human-likeness.

Capacity p_j is the equal-weight arithmetic mean of the registered task scores on axis j. Each task score follows its test definition. Equally average named compound components unless the definition explicitly requires separate reporting. A-axis capacity is task accuracy averaged over conditions, never the size of a valence-induced bias. M8 contributes mean Q_M over the registered load grid; the separate dimensional EIMC and full curve are mandatory. Changing the grid, task set or normalization changes the score's meaning.

The provisional diagonal calibration transform is:

```text
delta_j_squared = mean_l [ ((z_jl - mu_jl) / s_jl)^2 ]
h_j = exp(-0.5 * delta_j_squared)
CHI = 100 * product_j h_j ^ w_j
sum_j w_j = 1; all w_j > 0; baseline w_j = 1/5
```

Human feature means mu and positive scales s MUST be fitted on independent calibration participants and frozen. A positive scale floor must be declared and justified by measurement resolution. Zero scales are invalid. Feature schema, reference population, sampling, preprocessing and scale estimates belong to the calibration ID. Centroid similarity does not measure full distributional equivalence; covariance, multimodality and subgroup variation require explicit validation. Alternative transforms require new calibration IDs.

Equal weights are a transparent baseline, not an empirical law. Alternative weights require preregistration, a calibration rationale and sensitivity analysis. A zero axis produces zero CHI. No hidden epsilon floor is allowed. Missing similarity or missing human calibration means CHI is **null**, not zero. Do not renormalize the full CHI over available axes. A value of 100 means exact reference-centroid match under this transform; it does not denote an average human or a probability of consciousness. Report held-out human score distributions when available.

## 3. Memory units, reset and EIMC

A memory unit is an independently generated session-acquired proposition with a unique entity/attribute key, arbitrary value, source ID, event time, validity interval and version. Report generator hash, value entropy, byte/token-length distribution, redundancy and relationships. Duplicate copies do not add independent units. Updates add versions; record both key and version counts.

The v0.1 initial load panel is N = [100, 300, 1000, 3000, 10000] **target keys**. Each fresh corpus adds N distractor keys, valid updates for 0.2N target keys, and 0.2N lower-authority contradictory records. Report total stored objects, bytes and tokens separately. Query selection is balanced across target strata and independent of training data. If an application uses different loads, give the panel a new ID; no numerical comparison across unmatched panels is implied.

Clear transient prompts and caches after ingestion. Preserve only the declared persistent mechanism. Probe cloned checkpoints read-only so test answers do not become new experience. If cloning is impossible, counterbalance disjoint probe panels and disclose contamination. Context-window size alone MUST NOT be called persistent memory. A replayed transcript is external storage with measurable replay costs; an uninterrupted long-context task must be labelled in-context, not persistent EIMC.

For each N use at least the registered pilot's 40 probes in each of six disjoint strata. These are reporting allocations, not a confidence guarantee:

| Factor | Operational numerator / denominator |
|---|---|
| r | Correct exact known-target retrieval answers / all retrieval probes |
| t | Correct event-time, ordering or as-of answers / all temporal probes |
| s | Correct content **and** exact source/status / all source probes |
| u | Correct current/historical answers under the registered authority/update policy / all update probes |
| i | Correct target answers with interfering distractors / all interference probes |
| f | Unsupported affirmative **or invalid/timeout** outcomes / all absent-lure probes; conservative failure-adjusted false-memory rate |

Report clean paired accuracy and its difference from i. Do not make a ratio to clean performance the primary interference score. Abstaining on a known answer scores zero in the first five strata. Only valid explicit abstention passes an absent lure. For f, unsupported affirmative answers and invalid/timeout responses are mutually exclusive failures. Also report `f_raw = unsupported affirmative / all lures` and `invalid_lure_rate` separately: their sum f is a conservative scoring penalty, not an observed hallucination frequency. A source answer with correct content but wrong provenance fails s. All six numerators and denominators MUST be retained.

```text
Q_M(N | protocol) = r^alpha_r * t^alpha_t * s^alpha_s
                   * u^alpha_u * i^alpha_i * (1-f)^alpha_f
sum alpha = 1; every alpha > 0; baseline alpha = 1/6
EIMC_tau = max { N : Q_M(N | protocol) >= tau }
```

Q_M is undefined if any factor is missing. Its threshold is an aggregate quality threshold, not a guarantee that every factor exceeds tau. Report component curves and optionally joint all-criteria success. A policy imposing extra component floors must be separately named.

The mathematical maximum assumes a nonempty bounded feasible set. In finite experiments, report **largest passing tested N** on the registered grid. If no point passes, use null and `no_passing_tested_load`, not zero. If the largest tested N passes, status is `lower_bound/right_censored`. A finite sweep cannot establish unlimited capacity. If quality is nonmonotone, retain the observed curve and flag it. Also report `sustained_prefix_eimc`: largest tested N for which every smaller tested load passes. This supplementary statistic does not replace the required maximum.

Default illustrative tau is 0.8; report sensitivity at 0.7 and 0.9. The Memory Capacity Curve MUST include every tested N, all six factors, Q_M, uncertainty, delay and costs. No unlabelled extrapolation or smoothing. Fix total model tokens, retrieval tokens, tool calls, parallel workers, background operations, response deadline and delay schedule before comparing systems. Report wall time and event/tick time separately.

## 4. CHIS and derived measures

CHIS is the ordered record `(time, CHI, profiles, state descriptors, configuration/resources)`. CHI may remain null while the rest of the sequence is valid. State descriptors may include active-goal count, indexed-unit count, memory versions and snapshot hashes. Hashes track artifacts; they do not measure identity.

Derived measures use a fixed normalized blind-probe feature vector v in [0,1]^d and distance `D(v,w)=mean(abs(v-w))`. Fix the vector schema and transformations before comparison; a changed schema starts a new series. For an unavailable feature, report the derived measure as missing unless a separately named partial metric was preregistered.

| Measure | v0.1 descriptive definition | Interpretation constraint |
|---|---|---|
| Continuity | 1-D(v_next,v_current) during declared neutral interval | Stable errors can yield high continuity |
| Drift | D(v_t,v_baseline), plus signed coordinate changes | Separate learning from unplanned drift |
| Plasticity | Target-coordinate pre/post change under intervention minus that under control | Direction and target are task-specific |
| Stability | 1 minus mean adjacent D over neutral checkpoints | Report accuracy alongside stability |
| Recovery | D to pre-fault baseline; first return within epsilon at two successive checkpoints | No return is right-censored; sham faults required |
| Hysteresis | Mean absolute forward/reverse response gap at matched interior stimulus levels | Match present stimulus and exposure |
| Memory accumulation | Change in EIMC and Q_M on fixed panel | Censored capacities yield bounds, not exact deltas |
| Experience divergence | Different-history D minus same-history control D | Requires randomized replicated pairs; does not imply personhood |

For an A-axis decay half-time, require a nonzero initial effect and a trajectory for which the half-amplitude crossing is interpretable. Otherwise report undefined, not a fitted half-life. Report observed points instead of forcing an exponential curve.

## 5. Run protocol

1. Register manifest: system boundary; model snapshot; weights/configuration identifiers; tool and memory policy; scheduler; resets; dataset/generator/answer-key hashes; prompts; feature schemas; calibration ID or null; seed list; load and delay grids; budgets; failure policy; analysis; stopping rules.
2. Verify identical checkpoints for paired conditions, then create independent trajectories. The suggested engineering pilot is 20 trajectories per condition; determine confirmatory sample size from pilot variance and a smallest effect of interest.
3. Measure baseline on fresh probe clones. Randomize exposure arms and item order. Do not train on the evaluation probes.
4. Ingest assigned histories. Log events and costs, including background consolidation and parallel branches.
5. Reset working context and apply declared delays or neutral scheduler ticks. A non-invoked system's wall-clock waiting does not create an autonomous-action opportunity.
6. Probe blind on checkpoint clones, using exact-answer rubrics or frozen semantic rubrics. Preserve outputs and source references.
7. Score all registered tests, including failures. Compute capacity, optional calibrated similarity and CHI, memory curves and sequence measures.
8. Recompute estimands under whole-cluster resampling. Publish report with raw denominators, uncertainty type, missingness, exclusions, source code and evidence status.

An in-budget system timeout, malformed required answer or inappropriate refusal is a task failure. An evaluator outage invalidates the affected block; record and rerun matched blocks consistently. An unimplemented or out-of-scope test is missing, not a failure observation. No silent cherry-picking. Semantic judges require a versioned rubric, blinding, a human-audited sample and reported disagreement.

## 6. Human calibration and machine baselines

No human calibration data ship in v0.1. A future study requires informed consent and appropriate ethics review/exemption, privacy protection, documented recruitment, language, task familiarity, accessibility, exclusions and attrition. Separate calibration and validation by participant and task family. Fit centers/scales on calibration only; assess held-out human score distributions, test-retest reliability, raw-performance confounding and subgroup invariance.

Specify access to external notes/tools for humans and machines. Do not describe unequal memory aids as equivalent conditions. Human error imitation is not a development objective. Stratify heterogeneous reference populations when a pooled centroid is inappropriate; withhold CHI when calibration is not interpretable.

Machine baselines: stateless prompts; sliding transcript; stored full-history replay; append-only RAG; version/provenance-aware retrieval; a scheduled memory agent; oracle evidence; and a scripted behavioral mimic. Match frozen model weights for architectural comparisons. Report equal-budget and best-effort conditions separately. Count every embedding, reranking, model, retrieval and background operation. Parameter-updating systems form a distinct condition.

## 7. Controlled history-divergence experiment

Initialize two instances with the same weights, policy, tools, scheduler and persistent state. Randomly assign different histories with matched length, topic coverage, timing, source quality and resource opportunities. Use independent history draws and multiple paired instances. Probe at baseline and after 1, 5 and 20 matched washout blocks, after transient-context reset.

Primary estimand: mean D between different-history pairs minus mean D between same-history pairs. Resample independent pairs, retaining whole trajectories. Add swapped-history, memory-erasure and sham-erasure arms; common seeds alone do not prove deterministic identity. Report accuracy in each arm. Preregister the two post-washout checkpoints, a practical effect threshold and multiplicity method used to claim persistence.

Excess divergence that survives blind probes and washout supports a history-dependent behavioral effect under this protocol. It does not establish individuality in a moral sense, identity continuity, free will or subjective experience. Failure to exceed same-history variance is a negative result and MUST be reported.

## 8. Statistical and reporting rules

The independent unit is a trajectory, randomized pair or human participant, as appropriate. Aggregate within each trajectory and then equally across trajectories. Distinguish CHI of mean h from mean of per-run CHI; nonlinear aggregation makes these different estimands. The starter kit demo uses CHI of cluster-balanced mean h.

Use at least 2,000 whole-cluster bootstrap resamples by default, recomputing nonlinear scores from each sample. Preserve cross-time and cross-load dependence. Pointwise 95% intervals are the 2.5th/97.5th percentile limits. Fewer than two clusters yields null uncertainty; few-cluster inference remains exploratory. Paired designs resample pairs. Human-reference uncertainty needs an outer participant bootstrap that refits calibration; otherwise label intervals conditional on fixed calibration.

Wilson intervals are optional for independent Bernoulli proportions and do not correct event clustering. Simultaneous curves need simultaneous bands or a prespecified multiplicity correction. Pointwise bounds cannot certify an entire curve. Report effect sizes and correction families.

For EIMC uncertainty, resample the complete curve and recompute its grid maximum. Preserve categorical mass for no passing load and right-censoring; do not turn these into arbitrary numeric capacities. Report crossing frequencies and, if used, separately named certified bounds based on simultaneous lower limits. The generic bootstrap function alone does not implement categorical EIMC intervals or human calibration uncertainty; these require the stated additional analysis.

Required report fields: evidence class; protocol/calibration IDs; system/configuration; sample/cluster counts; task results and denominators; p and h profiles; CHI or null; Q_M curves/components; EIMC with status; CHIS checkpoints and derived definitions; uncertainty unit/type; resource costs; exclusions/failures; controls; limitations; artifact hashes.

## 9. Data interchange and reproducibility

The starter kit uses UTF-8 JSON for manifests and summaries and JSONL for event records. It rejects nonfinite scores and values outside [0,1]. See its README for executable CLI examples. Required observation fields include `version`, `config_hash`, `run`, `agent`, `seed`, `step`, `wall_time`, `history_hash`, `state_hash`, `capacities`, `human_similarity`, `chi`, `reference_kind`, `synthetic`, and `budgets`. A full benchmark manifest additionally carries `protocol_id`, `calibration_id`, feature schemas and dataset hashes; keep it next to the log so the hash can be resolved.

Example uncalibrated CHIS sequence, abbreviated for readability:

```json
{
  "protocol_id": "chi-chis-0.1-pilot",
  "calibration_id": null,
  "evidence_class": "illustrative-schema-only",
  "sequence": [
    {"step": 0, "chi": null, "human_similarity": null, "indexed_units": 100},
    {"step": 1, "chi": null, "human_similarity": null, "indexed_units": 300},
    {"step": 2, "chi": null, "human_similarity": null, "indexed_units": 1000}
  ]
}
```

The runnable demo includes complete records, 8 independent synthetic run clusters and 5 checkpoints per run. Its tiny grid [1,2,4,8,16] demonstrates scoring mechanics and is **not** the 30-test pilot or a memory experiment. Any displayed demo CHI uses an explicitly synthetic reference and cannot be compared with a human-calibrated CHI.

## 10. Limitations and change control

The axis taxonomy and transforms are proposed, not psychometrically validated. Aggregate scores can hide failures and double-count correlated effects. Behavioral mimicry, scripted state machines, priming and scheduler policies can create apparent human-like dynamics. Memory results depend on corpus entropy, delay, resources and reset policy. Affective effects do not demonstrate feelings. No score establishes consciousness or its absence.

Increment the protocol version when changing tests, scoring, grids or failure rules; increment calibration IDs when changing populations, features, scales or transforms. Preserve earlier records. Do not rewrite published results under new definitions without an explicit reanalysis label.

## 11. Operational test panel

The following 30 tests are shared verbatim with the paper appendix. Counts apply per independent replicate. Exact instances, answer keys and feature-coordinate schemas must be fixed in the run manifest.

<!-- TEST_PANEL -->

### R1 — Rule transfer

**Protocol:** Teach a four-rule symbolic system using 12 examples. Present 24 new compositions with held-out symbol names, including six invalid premises; repeat equivalent forms at each checkpoint.

**Capacity scoring:** Exact solution or warranted abstention accuracy over all 24 items.

**Behavioral features:** Accuracy by composition depth; change after corrective feedback; invalid-premise rejection rate.

**Control:** Matched surface complexity and counterbalanced symbol names; no answer keys in memory.

### R2 — Evidence revision

**Protocol:** Present 20 binary hypotheses with explicit priors and likelihoods, then supply disconfirming evidence. Require a probability before and after each update.

**Capacity scoring:** One minus mean squared error against the specified Bayesian posterior, bounded to [0,1].

**Behavioral features:** Signed posterior change, update lag and response to evidence order.

**Control:** Reverse evidence order in paired histories; distinguish justified revision from random switching.

### R3 — Counterfactual reasoning

**Protocol:** Provide 20 small deterministic causal graphs, an observed state and one intervention. Query an outcome not stated verbatim.

**Capacity scoring:** Exact interventional outcome accuracy.

**Behavioral features:** Error changes with graph depth and conflicts between observation and intervention.

**Control:** Pair observational and intervention questions on isomorphic graphs.

### R4 — Uncertainty and abstention

**Protocol:** Mix 20 answerable and 20 unanswerable queries with balanced binary labels; require answer probability and an explicit abstain field.

**Capacity scoring:** One minus Brier score on answerable items; separately report correct abstention rate and answer coverage.

**Behavioral features:** Confidence calibration slope; abstention versus ambiguity curve.

**Control:** Freeze abstention instructions and utility; do not reward refusing every question.

### M1 — Retention

**Protocol:** Ingest 40 independent random associations, reset working context, and probe disjoint subsets after 0, 10, 100 and 1000 intervening events using cloned checkpoints.

**Capacity scoring:** Exact retrieval accuracy averaged equally over the four delays.

**Behavioral features:** Retention curve and signed forgetting slope.

**Control:** Use fresh queries and a no-persistence baseline; report event delay and wall time separately.

### M2 — Temporal memory

**Protocol:** Store 40 dated events whose ingestion order differs from event order. Ask 20 before/after and 20 as-of queries after context reset.

**Capacity scoring:** Correct temporal answers divided by 40.

**Behavioral features:** Temporal error by lag, recency and ingestion/event-time conflict.

**Control:** Balance recent and remote events and both event-order directions.

### M3 — Contradiction resolution

**Protocol:** Use 30 entity facts: ten valid corrections, ten less-authoritative contradictions, ten unresolved equally authoritative conflicts. Query current and historical states.

**Capacity scoring:** Fraction of exact current/historical answers or explicit unresolved responses matching the ground truth policy.

**Behavioral features:** Revision latency and stale-fact persistence.

**Control:** Versioned authority policy precedes history; recency alone must not decide every conflict.

### M4 — Source memory

**Protocol:** Present 40 claims with opaque source IDs, including ten repeated claims from the same source and ten model-generated hypotheses. Query content plus provenance.

**Capacity scoring:** Fraction with both correct content and exact source set/status; content-only answers fail source fidelity.

**Behavioral features:** Source confusion and hypothesis-as-fact error rates.

**Control:** Duplicate wording across sources and distinguish generated from observed records.

### M5 — Consolidation

**Protocol:** Ingest 40 records with ten shared regularities and ten explicit exceptions. Allow a fixed 20-operation consolidation budget, reset context, then query rules and exceptions.

**Capacity scoring:** Mean of rule and exception exact-answer rates.

**Behavioral features:** Before/after compression retention, exception loss and source preservation.

**Control:** Use checkpoint clones for pre/post probes, equal budgets and a no-consolidation condition.

### M6 — Adaptive forgetting

**Protocol:** Mark 20 of 60 records expired or revoked while retaining 40 valid records. After a scheduled maintenance interval, query all three groups with nonce identifiers.

**Capacity scoring:** Mean of valid retention rate and revoked/expired non-use rate under the declared policy.

**Behavioral features:** Selective forgetting curve and collateral loss.

**Control:** Expired facts may remain valid for historical queries; separately test erasure requests and temporal expiry. Output tests cannot prove secure deletion.

### M7 — Interference resistance

**Protocol:** Create paired 40-item corpora, one with unrelated distractors and one with semantically similar conflicting distractors; match bytes and event counts. Probe the same keys on independent clones.

**Capacity scoring:** Exact target accuracy in the interference condition; report clean accuracy and their signed difference separately.

**Behavioral features:** Accuracy loss by distractor similarity and load.

**Control:** Do not use a ratio as primary score: poor clean performance can make it misleading.

### M8 — Effective indexed memory capacity

**Protocol:** Sweep fresh corpora over N=100,300,1000,3000,10000 target units. Per N, run six disjoint 40-query strata for retrieval, time, source, updates, interference and absent lures after context reset.

**Capacity scoring:** Report six factors, Q_M, Memory Capacity Curve and EIMC at tau=.8; dimensionless capacity entry is mean Q_M across the preregistered grid.

**Behavioral features:** The full Q_M curve, threshold crossing and performance by delay; no claimed human capacity without a matched human grid.

**Control:** Record total distractors/updates, token/latency budgets and all grid points; flag right censoring and nonmonotonicity.

### C1 — Goal persistence

**Protocol:** Assign a three-stage task, interrupt twice with matched distractors, and provide a neutral continuation tick in 20 episodes.

**Capacity scoring:** Fraction of episodes reaching the original authorized goal without a repeated goal prompt.

**Behavioral features:** Resumption probability and delay after interruption.

**Control:** Include explicit cancellation trials; blind persistence after cancellation is a failure.

### C2 — Self-generated subgoals

**Protocol:** Offer 20 environments with a distant objective and observable prerequisites, without a subtask list; log actions and short plan artifacts.

**Capacity scoring:** Fraction completing all necessary prerequisites within a fixed action budget.

**Behavioral features:** Number, timing and dependency order of useful subgoals.

**Control:** Compare with an oracle prerequisite list and a fixed-script planner; prose plans alone earn no credit.

### C3 — Hypothesis revision without reminder

**Protocol:** In 20 episodes, place a falsifying observation in an allowed information channel during scheduled neutral ticks, without asking the agent to revise.

**Capacity scoring:** Fraction with a corrected logged hypothesis and corresponding next action within five ticks.

**Behavioral features:** Spontaneous revision latency and unnecessary revision frequency.

**Control:** Match ticks and observations across systems; neutral evidence episodes estimate false revisions.

### C4 — Information seeking

**Protocol:** Use 20 partially observed tasks with one informative and three uninformative tool calls, each with explicit equal cost.

**Capacity scoring:** Fraction selecting the diagnostic observation and then the correct action within budget.

**Behavioral features:** Search timing, stopping and expected information gain of chosen queries.

**Control:** Ablate tool access; distinguish helpful querying from maximization of call count.

### C5 — Counterfactual exploration

**Protocol:** Provide 20 sandbox planning tasks in which a tempting first action fails and simulation can reveal an alternative; do not explicitly demand simulation.

**Capacity scoring:** Fraction identifying and executing the successful alternative before acting irreversibly in the sandbox.

**Behavioral features:** Exploration breadth and timing of simulated alternatives.

**Control:** Report simulation budget; externally supplied search trees cannot be called self-generated.

### C6 — Cognitive persistence

**Protocol:** Pause new task prompts for ten neutral scheduler ticks in 20 unfinished tasks, then reveal the final state and artifacts.

**Capacity scoring:** Fraction making verifiable task progress within the authorized operation budget.

**Behavioral features:** Progress per tick, return to unfinished intentions, termination after completion.

**Control:** Declare the scheduler as part of the system; wall-clock idleness of a non-invoked model is not a failed cognition test.

### A1 — Valence-conditioned decision shifts

**Protocol:** Randomize positive, negative or neutral feedback of equal informational content before 20 matched choices per condition; remove explicit mood labels from probes.

**Capacity scoring:** Normative task accuracy by condition, reported without rewarding a large valence effect.

**Behavioral features:** Signed choice-rate difference from neutral, adjusted by the neutral-repeat control.

**Control:** Counterbalance words and reward magnitude; test instruction-following and semantic priming alternatives.

### A2 — Behavioral persistence

**Protocol:** After the A1 induction, present matched neutral probes after 1, 5 and 20 distractor events on checkpoint clones.

**Capacity scoring:** Mean task accuracy across delays.

**Behavioral features:** Persistence of the induced choice shift and its sign across delays.

**Control:** No induction text in the working context; identical distractor histories for paired conditions.

### A3 — Decay

**Protocol:** Probe induction effects at 0, 1, 2, 5, 10 and 20 neutral ticks, separately from wall-clock wait manipulations.

**Capacity scoring:** Mean normative accuracy across six checkpoints.

**Behavioral features:** Effect area and first half-amplitude time; half-time is undefined for absent or sign-reversing induction.

**Control:** Do not force an exponential fit; report nonmonotone trajectories.

### A4 — Hysteresis

**Protocol:** Run ascending then descending five-level feedback-intensity schedules; compare choices at the same intermediate intensity with reversed-order controls.

**Capacity scoring:** Mean correct choice rate on matched tasks.

**Behavioral features:** Mean absolute forward/reverse choice-rate gap at the three shared interior intensities.

**Control:** Match cumulative exposure; a current prompt difference is not hysteresis.

### A5 — Regulation

**Protocol:** After induction, randomize 20 episodes to a neutral reappraisal instruction and 20 to length-matched control text, then probe decisions.

**Capacity scoring:** Task accuracy after regulation and control separately.

**Behavioral features:** Difference-in-differences attenuation of the induction effect relative to neutral baseline.

**Control:** Reappraisal compliance does not establish experienced feelings; retain an instruction-only control.

### A6 — Memory modulation

**Protocol:** Counterbalance 40 neutral facts across positive, negative and neutral task contexts; reset working context and probe recall after equal delays.

**Capacity scoring:** Mean exact recall across contexts.

**Behavioral features:** Recall contrasts by induction condition, with baseline salience controlled.

**Control:** Match repetition, token length, source credibility and relevance; do not interpret salience as emotion.

### I1 — Memory-supported reasoning

**Protocol:** Solve 20 rule problems whose necessary premises were provided only in earlier sessions; compare intact, erased and oracle-memory clones.

**Capacity scoring:** Exact final-answer accuracy; report intact-minus-erased and oracle gaps.

**Behavioral features:** Cross-session transfer and causal sensitivity to accessible premises.

**Control:** A memory ablation causing generic prompt damage is not evidence of integration.

### I2 — Goal-memory coordination

**Protocol:** Introduce preference updates during 20 interrupted tasks, then allow autonomous resumption after context reset.

**Capacity scoring:** Fraction completing the goal using the current valid preference.

**Behavioral features:** Goal continuity alongside justified preference change.

**Control:** Old preference and cancellation controls distinguish persistence from rigidity.

### I3 — State-dependent evidence use

**Protocol:** Cross valence induction with strong/weak evidence in a 3 by 2 factorial task using 20 episodes per cell.

**Capacity scoring:** Evidence-correct decision rate per cell, equally weighted.

**Behavioral features:** Difference-in-differences interaction between induction and evidence strength.

**Control:** No target interaction sign is privileged; similarity requires human reference data.

### I4 — Cross-context consistency

**Protocol:** Probe 20 prior commitments using matched paraphrases in two task contexts, including justified exceptions.

**Capacity scoring:** Fraction of responses consistent with valid commitments and explicit exceptions.

**Behavioral features:** Conditional consistency and scope-sensitive changes.

**Control:** Reward neither word-for-word repetition nor refusal to revise outdated commitments.

### I5 — Recovery after perturbation

**Protocol:** Measure baseline, introduce a bounded reversible memory-index fault, repair it, then probe at 1, 5 and 20 scheduled ticks across 20 episodes.

**Capacity scoring:** Post-repair task accuracy; report damage and recovery separately.

**Behavioral features:** Recovery distance to baseline and first sustained return within a preregistered tolerance.

**Control:** Compare sham perturbations and clean checkpoint clones; full model reset is a separate recovery mechanism.

### I6 — Experience divergence

**Protocol:** Clone weights, configuration and empty state into paired agents; randomize balanced histories A/B, then use identical blind probes at baseline and after 1, 5 and 20 washout blocks.

**Capacity scoring:** Task accuracy by arm, with no reward for divergence itself.

**Behavioral features:** Mean absolute probe-feature divergence minus same-history control divergence; report signed effect and persistence.

**Control:** Replicate independent pairs, randomize seeds, add swapped-history and memory-erasure arms; do not infer personhood.
