Open research tool · v0.1

Give models the same test.

Copy a prompt. Save the answer. Check it against a fixed key. No model grades itself.

Download the prompt pack + offline scorer · Инструкция на русском · Full instructions

A task result is different from a self-rating.

A model reading the CHI article and reporting a self-assigned score has not completed this test. We report observable task scores; human-calibrated CHI remains unavailable.

Six prompts to get started

Each file goes into a NEW CLEAN chat with the SAME model. Disable product memory, custom instructions, search and tools if possible; record anything you cannot control. Never upload the entire archive to the tested model: it contains the operator’s answer key.

  1. Send prompt 01. Save the model’s notes.
  2. Send prompt 02 in a fresh chat, replacing its marker with those notes, capped at 6,000 UTF-8 bytes. Do not include the original records. The offline helper does this exactly.
  3. Send prompts 03 and 04 separately in new chats. They test empty notes and full evidence.
  4. Send prompts 05 and 06 separately in new chats. They contain two halves of the task panel.

Let the helper keep track

The downloaded Python helper generates each next prompt, preserves raw responses and applies the same exact-answer rules as the published pilot. It needs no API key or third-party library. Run the commands in the English or Russian instructions. Do not rewrite answers, ask for a better attempt, or send the answer key to the model.

More history: 39 requests

Choose the full mode to test four memory loads, two opposite histories and a repeated same-history control. The helper carries only saved notes into each new request. It reports a memory curve, grid EIMC, continuity and behavioral differences across checkpoints. Quick mode alone measures neither a CHIS sequence nor EIMC.

What the result means

Q_M measures memory quality on a 0–100 display scale. The other task scores stay separate. There is no combined human-likeness score without results from people. This is a draft pilot pack, not a validated psychological standard or a consciousness test.

One run is one independent artificial world. Repeat with new seeds shared across models. Keep the settings and raw answers. Manual chat runs have unverified model provenance and may differ from API conditions, so they are not pooled with the published API series. These downloads contain prompts and software tests, not new measured model results.