# CPLOM publication evidence map

Reviewed 2026-09-20.

All seven publications listed in the local CPLOM publications index. This map separates claims from evidence; it does not assume the pilot confirms any article.

The pilot tests named models with the same bounded external notebook on artificial records and short scheduled histories. It does not test production CPLOM, provider-native memory or the full CHI/CHIS specification. CHI and human-similarity values remain null without human calibration.

Reviewed the seven local article or PDF texts and the pilot protocol. Publication titles follow the index. Operational source data, independent audits and control groups were not provided in this release. Related sources are background, not validation of CPLOM.

## [Beyond Intelligence: Introducing CHI and CHIS](https://cplom.ai/insights/chi-chis-humanlike-cognition.html)

**Main claim.** A system should be tested on how it remembers, changes and recovers over time. Comparing it with people requires a measured human reference.

**Evidence status.** Measurement proposal with a specification and calculation examples. The original release reports no measured CHI for a real model. This limited pilot cannot validate the full scale.

**This pilot can test.** Notebook memory quality across four loads; short history effects, persistence and correction; small reasoning, choice, framing and integration probes. Results may support or weaken these narrow hypotheses.

**This pilot cannot test.** Human likeness, a valid CHI ranking, the full 30-test panel, lasting individuality, autonomous thought, feelings or consciousness.

**Data still needed.** A defined human sample and matched conditions; held-out calibration; repeated full-panel measurements; longer histories; reliability and subgroup checks; tests of whether CHI adds value beyond existing benchmarks.

**Related research.** [longmemeval](https://arxiv.org/abs/2410.10813): Prior tests of updates, time, multi-session reasoning and abstention. [stratmem](https://aclanthology.org/2026.acl-long.1491/): Prior tests of memory use beyond factual recall.

## [From Control to Memory](https://cplom.ai/insights/from-control-to-memory.html)

**Main claim.** Useful long-term memory needs source records, dates, controlled updates and a way to search competing evidence.

**Evidence status.** Architecture proposal with small offline code examples. Those examples show control flow, not a complete memory system or measured advantage over alternatives.

**This pilot can test.** Whether a basic model-written notebook preserves correct values, sources, older facts and authoritative updates after a reset, and avoids unsupported answers.

**This pilot cannot test.** The proposed hierarchy, retrieval routing, competing evidence contexts, repeated challenge, deletion controls or production scale. The pilot notebook does not implement that architecture.

**Data still needed.** An implemented memory system; the same answering model and matched budgets across full-context, flat-retrieval and hierarchical alternatives; tests that remove each proposed feature; labeled evidence, update, access and deletion cases; cost and latency logs.

**Related research.** [rag](https://arxiv.org/abs/2005.11401): External retrieval is established prior work. [memgpt](https://arxiv.org/abs/2310.08560): Prior memory tiers and context-management architecture.

## [Why AI Needs Architecture to Become Infrastructure](https://cplom.ai/insights/why-ai-needs-to-be-arc.html)

**Main claim.** Reliable AI needs a process for checking constraints, uncertainty and alternatives before taking action.

**Evidence status.** Architecture argument informed by the author's deployment account. Broad reliability and economic claims are not established by a controlled comparison in this release.

**This pilot can test.** How the tested model-plus-notebook systems follow supplied constraints, revise choices and respond to missing evidence in small artificial tasks.

**This pilot cannot test.** Whether governance causes safer production decisions, whether shared-agent coordination improves operations, or whether total business costs fall. No governance architecture comparison is run.

**Data still needed.** Matched systems with and without specified checks; recorded actions and outcomes; defined failure and escalation rules; tests under disruption; workload, staffing, compute and total-cost records. A causal claim needs a credible control design.

## [From Correction to Adjudication](https://cplom.ai/insights/qpm-governance.html)

**Main claim.** Giving agents opposing roles and a formal approval process may catch errors that agreement alone misses.

**Evidence status.** Architecture proposal plus author-reported operational observations, not independently reproduced in this release. Reported route counts, correction rates, error reductions and savings lack the underlying operational data and control comparison here.

**This pilot can test.** Only separate model behavior on conflicting records and declared source rules. These are memory probes, not a test of QPM.

**This pilot cannot test.** QPM's advantage over voting or other review methods; the reported 37,500 routes with 22 corrections; roughly 20% residual-error reduction; labor savings or safe operation. A low manual-correction rate is not automatically a true error rate.

**Data still needed.** Route-level input, output and review logs; a clear error definition; independent review including uncorrected routes; baseline denominators and dates; QPM rules and versions; matched-compute voting and challenge controls; staffing, cost and service-outcome records.

## [From Reactive Optimization to Predictive Governance](https://cplom.ai/insights/predictive-governance.html)

**Main claim.** Coordinating routes, warehouses, staffing and computing may reduce unstable swings and extreme delays across the whole system.

**Evidence status.** Control framework plus author-reported operational observations, not independently reproduced in this release. Reported stability and transfer across deployments need underlying time-series data and a control design.

**This pilot can test.** Short changes in model choices after changed evidence, neutral notebook steps and a correction. This is a small behavioral analogy to tracking change, not a logistics stability test.

**This pilot cannot test.** Reduced delivery-time variance, fewer extreme service failures, weaker cross-region cascades, stable production control or transfer to other businesses.

**Data still needed.** Timestamped state, demand, weather, resource, intervention and service logs; fixed definitions of stability and extreme events; periods before and after rollout; comparison regions or phased rollout; uncertainty estimates and records from each claimed deployment.

## [From Metrics to Market Dynamics: Early Deployment Lessons](https://cplom.ai/insights/early-deployment.html)

**Main claim.** The author reports that coordinated predictive control increased delivery volume while reducing delivery time and active staffing.

**Evidence status.** Author-reported operational observations, not independently reproduced in this release. The reported 210% volume increase, 36% shorter average delivery time and 17% fewer active couriers are not results of this pilot.

**This pilot can test.** No direct deployment claim. Its memory and history tasks can help shape later model tests but contain no deliveries, dispatchers or warehouse operations.

**This pilot cannot test.** The reported throughput, time and workforce changes; dispatcher time savings; market effects; or whether CPLOM caused these changes rather than demand, staffing or other changes.

**Data still needed.** Original delivery, route, staffing and dispatcher-time records; exact dates and metric definitions; workload and service-quality measures; all concurrent process changes; a comparison group or credible rollout design; full costs and independent recalculation.

## [Cross-Layer Predictive Logistics Optimization Model](https://cplom.ai/CPLOM_White_Paper_v1.0.pdf)

**Main claim.** Combining forecasting, operational indices and rule checks across logistics layers may improve throughput and decision reliability.

**Evidence status.** Technical architecture description plus author-reported operational observations, not independently reproduced in this release. Equations describe the method; they do not verify the reported production effects.

**This pilot can test.** No direct production claim. It can measure only the declared notebook and short-history tasks on the selected public model endpoints.

**This pilot cannot test.** The reported 210% volume growth, 36% delivery-time reduction, 17% active-driver reduction, about 99.82% decision accuracy or about 800 ms computation time; production scalability or portability.

**Data still needed.** A versioned implementation and evaluation dataset; decision labels and denominators; raw timing and load traces; delivery and staffing records; comparison periods and systems; per-deployment results; tests separating forecasting, voting and rule checks. Agreement between repeated runs must be checked against actual correctness.
