Compile real-world Claude Code and Codex trajectories into verified, tradable post-training assets.
-
Updated
Aug 21, 2026 - Python
Compile real-world Claude Code and Codex trajectories into verified, tradable post-training assets.
Procedural data generators for verifiable reasoning, synthetic pretraining, post-training, evaluation, and RL.
Score the trustworthiness of outputs from any LLM in real-time
🐀 Fuzz your verifier before an RL agent does. Static + dynamic LLM security auditor to detect reward-hacking in RL post-training environments (OpenEnv, verifiers-spec, Gymnasium).
Open Arena: SLM verifiers across observability platforms, dataset and model catalogs, and value scenarios for LLM, agentic and harness evals.
An RL Enviorment for AES Inversion
Adversarial QA for LLM-RL environments: find out what reward an empty answer earns. Model-free, zero API cost.
A verifiers RLM environment for testing whether adaptive recursive search outperforms brittle manual RAG choreography on long synthetic corpora.
Verifiable RL environments for corporate law & governance — deterministic reward, no LLM judge, fully synthetic worlds.
An open reinforcement-learning (RL) environment that trains LLM agents to use the current fact, not the stale one — verifiable reward for temporal fact-currency, built on verifiers / prime-rl (GRPO, LoRA).
Reproducible verifier audits, datasheets, agreement metrics, and release gates
RLVR coding environment for training and evaluating LLM agents on automated code-fixing tasks.
Research proposal for verifier-gated on-policy distillation with explicit evaluation and claim boundaries
Verifiers hello world repo
Minimal verifier environments for testing agents.
RL environments for generating code whose correctness is checked by formal verification; first release: C + ACSL + Frama-C.
Deterministic SRT captioning environment for LLM evaluation and reinforcement learning.
A verifiers RL environment that trains models to propose novel, evidence-grounded, falsifiable hypotheses. Rewards novelty with accountability.
Three small demos on verifiable evaluation: measured retrieval accuracy, a validation loop that bounces bad work, and an end-to-end pipeline. No API key, no GPU.
Linear Algebra Done Right (Axler 4e) as an exactly-graded RL environment for the Prime Intellect Environments Hub. 186 tasks, 25 topics, binary reward from exact arithmetic over the rationals. No LLM judge, no network.
To associate your repository with the verifiers topic, visit your repo's landing page and select "manage topics."