The first open evaluation framework for AI continuity. 250 narrative tests, 1835 verification questions, 10 checkpoints. Benchmark for AI memory systems, stateful agents, and long-term context persistence. No LLM in the evaluation loop.
persistent-memory evaluation-framework memory-benchmark ai-agents continuity long-term-memory temporal-reasoning llm-evaluation ai-memory ai-benchmark context-engineering stateful-agents narrative-testing
-
Updated
Apr 21, 2026 - Python