# Bibliography and evidence audit

Checked 2026-09-20. This is a targeted primary-source audit, not a systematic literature review. The bibliography contains 16 entries. arXiv entries use the first public submission year, explicitly identify the preprint, and do not imply that the cited version received peer review. ACL entries use proceedings years. Full author lists were taken from arXiv citation metadata or ACL's exported BibTeX. The AI Magazine record supplies its journal metadata; Picard's PDF supplies title, author and report number, while MIT's catalog verifies 1995.

## Evidence and boundaries

| Key | Verified source and bibliographic date | What it supports | What it does not establish |
| --- | --- | --- | --- |
| `mmlu` | Hendrycks et al., *Measuring Massive Multitask Language Understanding*, [arXiv:2009.03300](https://arxiv.org/abs/2009.03300), first submitted 2020-09-07; ICLR 2021 noted in record. | A broad, subject-based evaluation of language-model knowledge and problem solving. | Persistent autobiographical memory, coherent identity, or a comprehensive definition of intelligence. Avoid describing all contemporary evaluation as MMLU-like. |
| `helm` | Liang et al., *Holistic Evaluation of Language Models*, [arXiv:2211.09110](https://arxiv.org/abs/2211.09110), 2022-11-16; TMLR 2023 noted in record. | Evaluation across scenarios and multiple desiderata; transparency and reporting tradeoffs. | That evaluation is confined to accuracy, or that a new scalar CHI score has established validity. |
| `lost` | Liu et al., *Lost in the Middle: How Language Models Use Long Contexts*, [arXiv:2307.03172](https://arxiv.org/abs/2307.03172), 2023-07-06. | Position-dependent use of information in the tested QA and retrieval settings; a reason to vary evidence placement. | A universal failure law for every 2026 model or proof that context length alone determines continuity. |
| `longbench` | Bai et al., *LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding*, [arXiv:2308.14508](https://arxiv.org/abs/2308.14508), 2023-08-28. | Existing evaluation across several long-context tasks and two languages. | Evaluation of an agent's full history of goals, commitments, recovery and identity across deployment. |
| `ruler` | Hsieh et al., *RULER: What's the Real Context Size of Your Long-Context Language Models?*, [arXiv:2404.06654](https://arxiv.org/abs/2404.06654), 2024-04-09. | Testing beyond simple needle retrieval, including complexity and length variation. | That nominal context capacity guarantees effective use, or that synthetic-task performance directly measures humanlike cognition. |
| `rag` | Lewis et al., *Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks*, [arXiv:2005.11401](https://arxiv.org/abs/2005.11401), 2020-05-22. | Combining retrieved nonparametric information with generation; a relevant retrieval baseline. | That all RAG systems are equivalent, or that retrieval by itself supplies persistent self-models and continuity. |
| `memgpt` | Packer et al., *MemGPT: Towards LLMs as Operating Systems*, [arXiv:2310.08560](https://arxiv.org/abs/2310.08560), 2023-10-12. | Hierarchical memory management and explicit context movement; prior architecture relevant to persistent agents. | Novelty of memory tiers in CHI, actual unlimited capacity, or guaranteed long-horizon reliability. |
| `generative` | Park et al., *Generative Agents: Interactive Simulacra of Human Behavior*, [arXiv:2304.03442](https://arxiv.org/abs/2304.03442), 2023-04-07. | Prior integration of experience records, retrieval, reflection and planning for simulated agents. | Consciousness, validated psychological identity, or indefinite continuity in open deployment. |
| `locomo` | Maharana et al., *Evaluating Very Long-Term Conversational Memory of LLM Agents*, [ACL 2024](https://aclanthology.org/2024.acl-long.747/), pp. 13851–13870; DOI 10.18653/v1/2024.acl-long.747. | Existing multi-session memory evaluation using QA, event summarization and multimodal dialogue tasks. | A claim that longitudinal conversation evaluation is absent from the literature, or that generated dialogues cover all real-world interaction. |
| `longmemeval` | Wu et al., *LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory*, [arXiv:2410.10813](https://arxiv.org/abs/2410.10813), 2024-10-14; ICLR 2025 noted in record. | Prior testing of extraction, multi-session reasoning, temporal reasoning, updates and abstention. | Novelty claims for these abilities alone. Broad continuity must add a distinct construct and demonstrate incremental validity. |
| `beam` | Tavakoli et al., *Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs*, [arXiv:2510.27246](https://arxiv.org/abs/2510.27246), first 2025-10-31; consulted v2 revised 2026-02-21. | BEAM's long-conversation memory evaluation and LIGHT's complementary memory components. | Calling this a first-published 2026 paper, or equating token count with real elapsed time, lifelong learning or identity. |
| `stratmem` | Wu, Wu, Zhu, Sha and Wang, *StratMem-Bench: Evaluating Strategic Memory Use in Virtual Character Conversation Beyond Factual Recall*, [ACL 2026](https://aclanthology.org/2026.acl-long.1491/), pp. 32309–32328; DOI 10.18653/v1/2026.acl-long.1491. | Strategic use of required, supportive and irrelevant memories in character dialogue; direct prior work beyond factual recall. | A claim that all memory benchmarks test only recall. Its character-dialogue scope does not validate CHI's wider proposed construct. |
| `standardmind` | Laird, Lebiere and Rosenbloom, *A Standard Model of the Mind: Toward a Common Computational Framework across Artificial Intelligence, Cognitive Science, Neuroscience, and Robotics*, [AI Magazine 38(4), 2017](https://ojs.aaai.org/aimagazine/index.php/aimagazine/article/view/2744/0), pp. 13–26. | Cognitive architecture as a principled source of design hypotheses; shared structure across established architectures. | A settled universal theory, or evidence that a particular LLM implementation reproduces the human mind. |
| `picard` | R. W. Picard, *Affective Computing*, [MIT Technical Report 321](https://www.cl.cam.ac.uk/~pr10/iui/picard95.pdf), 1995; [MIT catalog](https://hd.media.mit.edu/TechnicalReportsList.html) and [revised report](https://vismod.media.mit.edu/tech-reports/TR-321.pdf) corroborate year. | A historical rationale for computational treatment of affect and its role in interaction. | That affective variables entail felt emotion or consciousness. This 1995 technical report is distinct from the 1997 book. The Cambridge copy and revised MIT copy differ in layout/content; do not interchange page-specific citations. |
| `consciousness` | Butlin et al., *Consciousness in Artificial Intelligence: Insights from the Science of Consciousness*, [arXiv:2308.08708](https://arxiv.org/abs/2308.08708), 2023-08-17. | Theory-informed indicator properties and careful distinction between functional evidence and consciousness claims. | A validated consciousness test or evidence that a high CHI/CHIS score establishes experience or moral status. Its assessment of then-current systems is not a current 2026 survey. |
| `forgetting` | Kirkpatrick et al., *Overcoming catastrophic forgetting in neural networks*, [arXiv:1612.00796](https://arxiv.org/abs/1612.00796), first 2016-12-02; subsequent [PNAS article](https://doi.org/10.1073/pnas.1611835114), 2017. | Sequential task learning and retention as established research problems; motivation for measuring acquisition together with retention. | Equivalence between parameter-update catastrophic forgetting and failure to retrieve an external conversation memory. Use as conceptual context, not direct evidence for CHI mechanisms. |

## Implications for the proposal

1. Frame CHI/CHIS as a proposed operational synthesis and experimental program. None of these sources supplies CHI-specific observations, validated weights, thresholds, norms, reliability estimates or superiority results.
2. A defensible gap is narrower than “AI benchmarks ignore time/memory.” LoCoMo, LongMemEval, BEAM and StratMem-Bench already cover substantial portions of this territory. Specify which combination of persistent goals, commitments, temporal provenance, repair and cross-session behavior is not covered by the selected comparators, then test whether the proposed measures add explanatory value.
3. Separate model weights, prompt/context, retrieval store, system policies and action environment. A continuity result belongs to the evaluated system configuration, with its budgets and update rules, not automatically to the underlying model.
4. Do not infer human psychological constructs directly from implementation labels such as episodic, affective, reflection or self-model. Treat them as functional design terms unless separately validated.
5. Report component scores and uncertainty before adopting a composite. The proposed aggregate, its weights and any gating rules are hypotheses to preregister and sensitivity-test, not source-supported scientific facts.

## Verification scope

All requested source identities resolved to primary records. Exact titles and author lists were checked; arXiv dates came from submission histories/citation metadata. ACL BibTeX was fetched from the Anthology. AI Magazine metadata was checked against the publisher record after the DOI resolver failed in the browsing tool. The evidence mapping above is based primarily on abstracts and bibliographic records, with Picard's primary PDF inspected for identity; it is not a full reproduction or methods audit. No numerical performance claim from these papers has been independently reproduced. Any detailed experimental comparison should inspect and cite the relevant full-text methods, versions and conditions.
