Historical multi-model Backrooms experiment with configurable conversation templates and example transcripts.
-
Updated
Aug 14, 2026 - Python
Historical multi-model Backrooms experiment with configurable conversation templates and example transcripts.
Context-driven valuation bias and halo effects across six multimodal LLMs (companion study to Lee, 2026)
Local-first read-only tools for agent context, memory, and governance evaluation before agents act.
Intrinsic preferences of AI coding agents under underspecified prompts: Experiments across models (Claude, Gemini, GPT, etc)
Reproducible experiments about model behavior, with protocols, tool traces, counterexamples and uncertainty.
LLM 归因行为测试型评测基准:基于多情境任务比较模型对能动性、自由意志与责任的归因,并提供可复现运行、结构化计分与结果审计。
Этот репозиторий посвящен исследованию онтологических патологий в LLM-архитектурах. Я не ищу дыры в цензуре, я строю систему исследования и управления интеллектом, картографирую симуляционные побочные эффекты под давлением современных методов элаймента.
Public model-behavior evaluation harness focused on evidence selection, hypothesis formation, and debugging decisions
Longitudinal LLM behavior evals with matched controls for identity stability, contextual adaptation, and conversational continuity.
Instrument-gated evaluation engine for replaying known-outcome model behavior
Open-source LLM evaluation framework for multilingual and sociolinguistic robustness, model behavior, and controlled language variation.
Research-style evaluation of local LLMs focused on paired comparisons, uncertainty, benchmark stability, and trustworthy model-ranking conclusions.
SAE context and intervention experiments in Gemma 3 4B: methods, prediction baselines, results and reproducible evidence.
Reproducible LLM evaluation harness for failure analysis, mitigation experiments, and regression testing.
Pre-registered evaluation of sycophancy in frontier LLMs: matched-pair prompts, multi-turn pressure tests, blind human gold labels, an LLM judge validated at κ=0.89, and a system-prompt mitigation verified against true-premise controls.
Behavioral case study of author-age misattribution, miscalibrated confidence, and confession-without-correction across six frontier LLMs (N=3,225)
Replication package for the auditable dataset-lineage protocol
Public-facing repository for LLM evaluation, model behavior observation, drift and failure analysis.
To associate your repository with the model-behavior topic, visit your repo's landing page and select "manage topics."