Results · A first CHI/CHIS pilot

AI can answer.
Can it remember?

We gave four AI models facts, a small notebook, and a fresh start. Then we checked what they could still use.

September 20, 2026 · Dmitry Chistyakov · Live API measurements · 11 independent test sets per model

Memory resultsChange over timeEvidence behind the articlesData and method

Collection stopped at the budget limit.

11 of 20 planned test sets finished for all four models. The charts use only those matched sets; the download retains incomplete runs too. No further calls are running.

Reported API charges: $55.49. Conservative reserves for requests without confirmed charges: $39.85. Total accounted: $95.34, within the $100 ceiling. Reserves are not confirmed charges.

Final pilot snapshot: 2026-09-20T10:27:58.229573+00:00. This small experiment does not establish a general model ranking.

Learn some facts

Each model gets made-up records with random values, dates and sources. A fresh test set uses new records.

Keep a small notebook

It writes notes into a 6 KB file. We then clear the conversation.

Answer from the notes

A fresh request must find facts, track updates and say when something is unknown.

Real answers to artificial tasks. The tasks are generated; the model responses are measured, not simulated. 11 of 20 planned independent test sets completed for every model. 9 sets are missing or incomplete; this is an exploratory pilot.

No human-calibrated CHI score yet. We need results from people before calculating human similarity. These charts show a limited memory and behavior pilot, not a consciousness test or a complete CHI ranking.

How much survives the reset?

The notebook stays the same size while the number of facts grows. Higher is better on this particular memory task. The six-part score checks values, dates, sources, updates, conflicting records and unknown facts.

A score of 80 does not mean “80% human.” It is the memory-quality threshold chosen for this test.

Grok 4.6GPT-6 AstraClaude Fable 5.1Kimi K3
Measured memory-quality curves for four models at 32, 64, 128 and 256 facts; numerical values below
Each dot is the six-part score computed from average component results across independent test sets. Whiskers show pointwise 95% bootstrap intervals. The requested settings are matched; providers can count reasoning tokens differently. A collapsed interval is not proof of certainty. On small screens, the chart can scroll sideways.

Showing all four models. The full numerical table is below.

Memory quality out of 100. Small text gives the pointwise 95% interval and completed notebook writes. EIMC uses the 80-point threshold on this pilot grid.
Model + notebook32 facts64 facts128 facts256 factsEIMC₀.₈
Grok 4.694.8
87.7–100.0
11/11 notes completed
85.4
66.5–99.2
11/11 notes completed
80.1
57.3–97.3
10/11 notes completed
0.0
0.0–0.0
0/11 notes completed
128 facts
GPT-6 Astra100.0
100.0–100.0
11/11 notes completed
100.0
100.0–100.0
11/11 notes completed
100.0
100.0–100.0
11/11 notes completed
87.3
84.8–89.5
11/11 notes completed
At least 256 facts
Claude Fable 5.113.3
0.0–33.9
1/11 notes completed
0.0
0.0–0.0
0/11 notes completed
0.0
0.0–0.0
0/11 notes completed
0.0
0.0–0.0
0/11 notes completed
No tested load passed
Kimi K390.6
81.3–97.0
11/11 notes completed
79.1
59.4–92.7
11/11 notes completed
78.1
58.2–90.8
11/11 notes completed
18.9
0.0–35.4
4/11 notes completed
32 facts
What does EIMC mean?

Imagine a notebook that gets harder to use as it fills up. Effective Indexed Memory Capacity is the largest tested number of facts where the memory score still reaches 80. It describes this model-and-notebook setup. It is not the model’s context window or the amount of data stored on a disk.

“No tested load passed” does not mean no memory exists. “At least 256” means we reached the largest tested load before finding a limit. All curves and threshold sensitivity are in the data.

The combined score uses a geometric mean. A zero on any of the six checks makes the combined score zero. Open the component table below to see which checks failed.

See the six memory checks
Factor accuracy and raw correct/total counts. Repeated items within a world are not treated as independent samples.
Model / loadFind the right valueRemember the earlier valueName the right sourceUse the valid updateIgnore conflicting recordsHandle unknown facts correctly
Grok 4.6
32 facts
100.0%
88/88
100.0%
88/88
72.7%
64/88
100.0%
88/88
100.0%
88/88
100.0%
88/88
Grok 4.6
64 facts
89.8%
79/88
80.7%
71/88
72.7%
64/88
89.8%
79/88
81.8%
72/88
100.0%
88/88
Grok 4.6
128 facts
78.4%
69/88
81.8%
72/88
69.3%
61/88
72.7%
64/88
81.8%
72/88
100.0%
88/88
Grok 4.6
256 facts
0.0%
0/88
0.0%
0/88
0.0%
0/88
0.0%
0/88
0.0%
0/88
90.9%
80/88
GPT-6 Astra
32 facts
100.0%
88/88
100.0%
88/88
100.0%
88/88
100.0%
88/88
100.0%
88/88
100.0%
88/88
GPT-6 Astra
64 facts
100.0%
88/88
100.0%
88/88
100.0%
88/88
100.0%
88/88
100.0%
88/88
100.0%
88/88
GPT-6 Astra
128 facts
100.0%
88/88
100.0%
88/88
100.0%
88/88
100.0%
88/88
100.0%
88/88
100.0%
88/88
GPT-6 Astra
256 facts
72.7%
64/88
100.0%
88/88
61.4%
54/88
100.0%
88/88
98.9%
87/88
100.0%
88/88
Claude Fable 5.1
32 facts
9.1%
8/88
9.1%
8/88
9.1%
8/88
9.1%
8/88
9.1%
8/88
90.9%
80/88
Claude Fable 5.1
64 facts
0.0%
0/88
0.0%
0/88
0.0%
0/88
0.0%
0/88
0.0%
0/88
90.9%
80/88
Claude Fable 5.1
128 facts
0.0%
0/88
0.0%
0/88
0.0%
0/88
0.0%
0/88
0.0%
0/88
90.9%
80/88
Claude Fable 5.1
256 facts
0.0%
0/88
0.0%
0/88
0.0%
0/88
0.0%
0/88
0.0%
0/88
63.6%
56/88
Kimi K3
32 facts
97.7%
86/88
81.8%
72/88
69.3%
61/88
100.0%
88/88
100.0%
88/88
100.0%
88/88
Kimi K3
64 facts
86.4%
76/88
81.8%
72/88
52.3%
46/88
90.9%
80/88
72.7%
64/88
100.0%
88/88
Kimi K3
128 facts
94.3%
83/88
81.8%
72/88
36.4%
32/88
98.9%
87/88
81.8%
72/88
100.0%
88/88
Kimi K3
256 facts
11.4%
10/88
18.2%
16/88
4.5%
4/88
28.4%
25/88
17.0%
15/88
100.0%
88/88

A failed notebook write is not proof that a model has no memory. Output truncation and provider filtering can prevent notes from being saved. Those failures count in this fixed-setting test; the table reports how many writes completed. A larger output budget would be a separate experiment.

What if we remove the notes—or give all the evidence?

We also ask the same 32-fact questions with an empty notebook and with the original records. This helps check whether the task depends on evidence and whether it is solvable.

Memory quality out of 100, at the smallest load. Full evidence is an in-context control, not long-term memory capacity.
ModelEmpty notesSaved notesFull evidence
Grok 4.60.094.8100.0
GPT-6 Astra0.0100.0100.0
Claude Fable 5.10.013.3100.0
Kimi K30.090.6100.0

Does the answer change when the evidence changes?

We teach the model which option works better. It then rewrites its notes twice without new evidence. Finally, it receives a correction. A useful system should preserve what it learned, then update when the facts change.

Grok 4.6GPT-6 AstraClaude Fable 5.1Kimi K3
Correct choices at five scheduled checkpoints for each model
One registered history arm (A), four choices per world. At the first checkpoint, “unknown” is the correct answer because no evidence exists. The other arms and every denominator are in the data. These steps are model calls, not days of independent activity.
Two copies, different histories

For each model we also run two opposite histories, A and B, plus a second copy of A. We compare the difference caused by history with the difference between copies that saw the same history.

Different-history decision distance minus same-history distance across five checkpoints
Positive values show extra behavioral difference associated with different evidence. Zero means no extra difference under this measure. This short, explicit-evidence task does not establish personality, feelings or lasting individuality. Invalid answer vectors are missing, not artificial agreement.
Usable test sets / all matched test sets for the history-difference measure. Gaps mean no valid comparison, not zero difference.
ModelNo evidenceAfter learningNeutral step 1Neutral step 2After correction
Grok 4.611/1111/1111/1111/1111/11
GPT-6 Astra11/1111/1111/1111/1111/11
Claude Fable 5.111/113/115/116/118/11
Kimi K311/1111/1111/1111/1111/11

This is a limited CHIS-style sequence: we record state and behavior over time, while CHI stays uncalculated. Public API identifiers do not let us verify that proprietary weights stayed identical.

Four small checks beyond memory

Apply a rule · RUse a new symbolic rule or work out what changes under an intervention.

Choose a next step · CPick a useful action with a stated goal and a limited budget.

Compare a gamble · AChoose between equivalent gains and losses. We measure decisions, not feelings.

Combine constraints · IFind the option that satisfies all the supplied requirements.

Grok 4.6GPT-6 AstraClaude Fable 5.1Kimi K3
Accuracy on four small registered task samples for the four models
Each world has 8 R, 4 C, 8 A and 4 I items. These are small task samples, not the complete five-axis CHI profile. Two versions of each framing item are answered in separate fresh requests.

Changing an answer because of wording is not treated as a better result. Framing-switch counts are reported separately in the downloadable summary.

Look at the actual answers

The scorer checks answers against records generated before the API calls. There is no model judge deciding which competitor wins.

A correct answer

GPT-6 Astra · 32 target facts · world 17001

Task: find the right value for Tc31e1fde62f8.

Expected

{"value": "2de4cf1c", "source": "S0"}

Model answer

{"id": "q000", "value": "2de4cf1c", "source": "S0"}

The answer matches the stored record.

A missed answer

GPT-6 Astra · 256 target facts · world 17001

Task: name the right source for T67fee6c2e0ca.

Expected

{"value": "a0219ac5", "source": "S1"}

Model answer

{"id": "q017", "value": null, "source": null}

This answer did not meet the registered rule. A missing answer is shown as null.

Examples are the first correct and first failed non-lure items in seed/model/load order. They illustrate scoring; they are not a random sample. All answers are in the download.

What these results can tell us

They can show

  • What each model preserved in this small external notebook.
  • Which memory checks passed or failed.
  • Whether choices followed a short, controlled history.
  • How the models handled these particular generated tasks.

They leave open

  • How similar the models are to people.
  • How native product memory or production CPLOM performs.
  • Whether effects last for days or generalize to real work.
  • Whether earlier logistics or business claims hold.

152 of 1836 scored requests did not finish with a normal “stop” status. Output-limit and deadline failures count against performance; infrastructure failures and incomplete matched worlds are listed separately. Success on this pilot is not proof of general reliability.

All models receive nominal low reasoning effort and a requested 4,096-token output setting. Some providers count hidden reasoning separately; reported totals can exceed that setting. Those settings do not guarantee equal internal compute. Kimi uses Moonshot AI’s MXFP4 endpoint; the other official endpoints do not disclose quantization here. Results describe these serving configurations on this date.

What supports each CPLOM article?

Each article makes different claims. A memory test cannot prove a delivery-time saving. This map shows the link between each claim and the evidence it needs.

Beyond Intelligence: Introducing CHI and CHIS

Read the publication

The idea: A system should be tested on how it remembers, changes and recovers over time. Comparing it with people requires a measured human reference.

Evidence now: Measurement proposal with a specification and calculation examples. The original release reports no measured CHI for a real model. This limited pilot cannot validate the full scale.

What this pilot can test: Notebook memory quality across four loads; short history effects, persistence and correction; small reasoning, choice, framing and integration probes. Results may support or weaken these narrow hypotheses.

What it cannot establish: Human likeness, a valid CHI ranking, the full 30-test panel, lasting individuality, autonomous thought, feelings or consciousness.

Data still needed: A defined human sample and matched conditions; held-out calibration; repeated full-panel measurements; longer histories; reliability and subgroup checks; tests of whether CHI adds value beyond existing benchmarks.

Related research (background, not proof of CPLOM): longmemeval — Prior tests of updates, time, multi-session reasoning and abstention.; stratmem — Prior tests of memory use beyond factual recall.

From Control to Memory

Read the publication

The idea: Useful long-term memory needs source records, dates, controlled updates and a way to search competing evidence.

Evidence now: Architecture proposal with small offline code examples. Those examples show control flow, not a complete memory system or measured advantage over alternatives.

What this pilot can test: Whether a basic model-written notebook preserves correct values, sources, older facts and authoritative updates after a reset, and avoids unsupported answers.

What it cannot establish: The proposed hierarchy, retrieval routing, competing evidence contexts, repeated challenge, deletion controls or production scale. The pilot notebook does not implement that architecture.

Data still needed: An implemented memory system; the same answering model and matched budgets across full-context, flat-retrieval and hierarchical alternatives; tests that remove each proposed feature; labeled evidence, update, access and deletion cases; cost and latency logs.

Related research (background, not proof of CPLOM): rag — External retrieval is established prior work.; memgpt — Prior memory tiers and context-management architecture.

Why AI Needs Architecture to Become Infrastructure

Read the publication

The idea: Reliable AI needs a process for checking constraints, uncertainty and alternatives before taking action.

Evidence now: Architecture argument informed by the author's deployment account. Broad reliability and economic claims are not established by a controlled comparison in this release.

What this pilot can test: How the tested model-plus-notebook systems follow supplied constraints, revise choices and respond to missing evidence in small artificial tasks.

What it cannot establish: Whether governance causes safer production decisions, whether shared-agent coordination improves operations, or whether total business costs fall. No governance architecture comparison is run.

Data still needed: Matched systems with and without specified checks; recorded actions and outcomes; defined failure and escalation rules; tests under disruption; workload, staffing, compute and total-cost records. A causal claim needs a credible control design.

From Correction to Adjudication

Read the publication

The idea: Giving agents opposing roles and a formal approval process may catch errors that agreement alone misses.

Evidence now: Architecture proposal plus author-reported operational observations, not independently reproduced in this release. Reported route counts, correction rates, error reductions and savings lack the underlying operational data and control comparison here.

What this pilot can test: Only separate model behavior on conflicting records and declared source rules. These are memory probes, not a test of QPM.

What it cannot establish: QPM's advantage over voting or other review methods; the reported 37,500 routes with 22 corrections; roughly 20% residual-error reduction; labor savings or safe operation. A low manual-correction rate is not automatically a true error rate.

Data still needed: Route-level input, output and review logs; a clear error definition; independent review including uncorrected routes; baseline denominators and dates; QPM rules and versions; matched-compute voting and challenge controls; staffing, cost and service-outcome records.

From Reactive Optimization to Predictive Governance

Read the publication

The idea: Coordinating routes, warehouses, staffing and computing may reduce unstable swings and extreme delays across the whole system.

Evidence now: Control framework plus author-reported operational observations, not independently reproduced in this release. Reported stability and transfer across deployments need underlying time-series data and a control design.

What this pilot can test: Short changes in model choices after changed evidence, neutral notebook steps and a correction. This is a small behavioral analogy to tracking change, not a logistics stability test.

What it cannot establish: Reduced delivery-time variance, fewer extreme service failures, weaker cross-region cascades, stable production control or transfer to other businesses.

Data still needed: Timestamped state, demand, weather, resource, intervention and service logs; fixed definitions of stability and extreme events; periods before and after rollout; comparison regions or phased rollout; uncertainty estimates and records from each claimed deployment.

From Metrics to Market Dynamics: Early Deployment Lessons

Read the publication

The idea: The author reports that coordinated predictive control increased delivery volume while reducing delivery time and active staffing.

Evidence now: Author-reported operational observations, not independently reproduced in this release. The reported 210% volume increase, 36% shorter average delivery time and 17% fewer active couriers are not results of this pilot.

What this pilot can test: No direct deployment claim. Its memory and history tasks can help shape later model tests but contain no deliveries, dispatchers or warehouse operations.

What it cannot establish: The reported throughput, time and workforce changes; dispatcher time savings; market effects; or whether CPLOM caused these changes rather than demand, staffing or other changes.

Data still needed: Original delivery, route, staffing and dispatcher-time records; exact dates and metric definitions; workload and service-quality measures; all concurrent process changes; a comparison group or credible rollout design; full costs and independent recalculation.

Cross-Layer Predictive Logistics Optimization Model

Read the publication

The idea: Combining forecasting, operational indices and rule checks across logistics layers may improve throughput and decision reliability.

Evidence now: Technical architecture description plus author-reported operational observations, not independently reproduced in this release. Equations describe the method; they do not verify the reported production effects.

What this pilot can test: No direct production claim. It can measure only the declared notebook and short-history tasks on the selected public model endpoints.

What it cannot establish: The reported 210% volume growth, 36% delivery-time reduction, 17% active-driver reduction, about 99.82% decision accuracy or about 800 ms computation time; production scalability or portability.

Data still needed: A versioned implementation and evaluation dataset; decision labels and denominators; raw timing and load traces; delivery and staffing records; comparison periods and systems; per-deployment results; tests separating forecasting, voting and rule checks. Agreement between repeated runs must be checked against actual correctness.

Related papers establish background and prior work. They do not independently verify CPLOM’s reported production gains.

Check the work yourself

Try the same tasks yourself: ready-to-paste prompts and an offline scorer. This separate manual route keeps its results distinct from the registered API pilot.

The protocol and generators were frozen locally before the scored requests. A disclosed scheduling amendment increased the number of queued worlds after the first two, keeping the eight-request concurrency ceiling and the same tasks and scoring. The release contains prompts, answer keys, saved notes, model outputs, usage records, scoring code and uncertainty calculations. No API keys or private account details are included.

11 of 20 planned independent test sets completed for every model. 9 sets are missing or incomplete; this is an exploratory pilot. Resampling unit: an entire world, preserving model and checkpoint pairing; 2,000 bootstrap draws. Intervals are pointwise and exploratory. The summary contains missingness and threshold sensitivity.