Open-source adversarial testing engine, SDK, and CLI for AI agents. Runs locally or against the Humanbound Platform.
-
Updated
Oct 2, 2026 - Python
Open-source adversarial testing engine, SDK, and CLI for AI agents. Runs locally or against the Humanbound Platform.
The Self-Evolving Agent Ecosystem — Trading agents that evolve through Darwinian selection and adversarial self-play
AI Robustness Evaluation System
Open-source test harness for AI agents. Stress-test production agents with adversarial multi-turn scenarios in CI
Toolkit for AI whitehats, internal red teams, llm bug bounty hunters and mlops - Adversarial testing for AI, LLMs, and Agents.
Open-source framework for building and testing LLM-powered applications: IRIS (single-agent orchestration), AETHER (declarative multi-agent systems), and AEGIS (adversarial security testing). Developed at MSU Denver's Community-Centered Computing (C3) Lab.
Your coding agent says it works. Gopnik tries to prove it doesn't — against the code, the delivered revision, and the checks themselves.
Opt-in Codex Skill and plugin for reliable coding: reactive failure recovery, elastic agent teams, task DAGs, one canonical writer, and assured delivery.
Production-ready Claude Code decision intelligence marketplace with adversarial routing, evidence-driven analysis, specialized agents, benchmarks, and automated validation.
Formal adversarial testing of LLM-generated industrial robot (URScript) code: ISO 10218-1:2025 safety + CWE security analysis, AST static watchdog, URSim runtime. Simulation-only.
Multi-perspective code review council for Claude Code. 3 advisors by default, 10 agents in deep mode (Opus + Codex). Evidence chains, adversarial self-test, dual-path verdict. Based on Karpathy's LLM Council.
Benchmark LLM jailbreak resilience across providers with standardized tests, adversarial mode, rich analytics, and a clean Web UI.
Autonomous flight simulator for AI agents — co-evolutionary adversarial red-teaming with 3D MAP-Elites quality-diversity. Discovers zero-days in LLM agents before production. MCP-native.
A deliberately weak support agent and the config to attack it locally with Humanbound
Official GitHub Actions for Humanbound — adversarial security testing for AI agents in CI.
AI safety evaluation framework testing LLM epistemic robustness under adversarial self-history manipulation
Context engineering toolkit for LLMs — pack, cache, debug, red-team, and orchestrate context windows. Council of Experts, adversarial testing, immune system, context compiler, drift detection, multi-agent entanglement. TypeScript + Python.
An automated red-teaming and reliability-auditing platform for AI agents - tests for prompt injection, tool hijacking and data exfiltration. Ships as an MCP server.
Adversarial testing for LLM applications. Calibrated prompt-injection and jailbreak scans with reproducible reports. Pip install, async-first, framework-agnostic.
To associate your repository with the adversarial-testing topic, visit your repo's landing page and select "manage topics."