A curated collection of publicly available resources on how technology and tech-savvy organizations around the world practice Site Reliability Engineering (SRE)
-
Updated
Nov 17, 2025 - JavaScript
A curated collection of publicly available resources on how technology and tech-savvy organizations around the world practice Site Reliability Engineering (SRE)
The open-source TypeScript framework for building AI workflows and agents. Designed for Claude Code describe what you want, Claude builds it, with all the best practices already in place.
Easily run integration tests for your backends
Portable, independent, web-based, simple streaming YouTube video queues and playlists for music videos, audiobooks, etc.
When the stakes are high, intelligence is only half the equation - reliability is the other
Examples of Error Handling and reaching High Reliability with vanilla JavaScript
A module to get the minimum usable engine(s)
Reliability toolkit for Google Flow — HAR forensics, fan-out analysis, eng brief. Not a bypass tool.
The open benchmark for measuring the reliability, reproducibility, and determinism of AI context systems. Transparent metrics, reproducible tests, vendor-neutral results. Context Trust Levels CTL 0-4.
Build AI agent workflows that survive crashes, never duplicate work, and don't make things up — one skill for Claude Code, Codex, Gemini CLI, Cursor, Windsurf/Devin, Hermes, and Copilot.
Simple Sloth SLO generator on the browser using Sloth as a Go library and WASM.
Robust React Programming
Uses esbuild to convert npm packages to ESM bundles for the browser.
DeepSeek Harness watchdog plugin: detects truly stalled agent turns (never killing in-progress tasks — in-flight operations are exempt), nudges/terminates only on real silence, records every event to JSONL with a loopback status route
Practical research, field notes, and reproducible experiments on AI agents, Claude & Codex workflows, and human-AI collaboration.
MCP server: declarative network-egress firewall for agent tools. Wraps @mukundakatta/agentguard.
Read-only Cloudflare usage watchdog that warns before Atlas Systems approaches configured Workers and KV usage ceilings.
MCP server: structured-output enforcer for any LLM. Wraps @mukundakatta/agentcast.
OpenClaw gateway reliability checks: restart preflight, diagnostics, watchdog, and synthetic stress probes for long-running agent systems.
Interactive mechanism demo for robot-policy reliability audits and objective failure scoring.
To associate your repository with the reliability topic, visit your repo's landing page and select "manage topics."