A compute-resource control plane. It is the authority on what compute exists, what it is capable of, whether it is healthy, whether a given model will fit on it, and where a workload should run.
As of v0.1.0 there is a minimal runtime: the computeconnect package serves both API layers
from one backend. Licensed Apache-2.0 (LICENSE). See
docs/STATUS.md — including its honest D2 re-evaluation — before proposing work.
python3 -m venv .venv && .venv/bin/pip install -e . # or: uv pip install -e .
.venv/bin/computeconnect serve --port 8090 --upstream http://127.0.0.1:8080--upstreamis an existing OpenAI-compatible llama.cpp server, consumed read-only — ComputeConnect never starts, stops, loads, or unloads it.- Port 8090 by default (on the reference host, 8080 is the engine and 8787 is BrainConnect's).
- Layer 1 (AgentConnect control plane):
GET /health,GET /models,GET /models/loaded,POST /route/estimate,POST /generate(streams; returnsX-Run-Id),POST /runs/{run_id}/cancel. - Layer 2 (OpenAI-compatible):
GET /v1/models,POST /v1/chat/completions. - Privacy is structural: no
privacy_tiermeans the most restrictive tier — cloud-class providers are filtered before placement and refusals are structured, never silent downgrades.
Tests: .venv/bin/python -m pytest (installs pytest via pip install -e .[dev]; real-engine
tests skip when no llama.cpp is reachable on :8080, and the two-engine subset also skips when no
second engine is reachable on :8091 via scripts/second_engine.sh). 154 tests total, 143 of them
fully offline; see docs/STATUS.md for the current pass count and the real-engine
caveat.
Every layer of the local-AI stack today assumes compute already exists and is already running: gateways route requests to endpoints somebody else started, engines load models onto hardware somebody else described, schedulers place pods on nodes somebody else registered. ComputeConnect is the layer underneath that answers what is actually out there, right now, and will this fit — and it is deliberately not an inference engine.
ComputeConnect does not perform inference. It never loads a tensor, never picks a
quantization, never implements an attention kernel. It knows that llama.cpp on the ARM box can
run a 30B MoE, and it knows how to ask it to; it does not know how.
It also does not own tasks, memory, tools, workflow engines, inference engines, or secrets managers. Those belong elsewhere, and the boundaries are drawn in docs/ARCHITECTURE.md.
ComputeConnect is the compute plane of the Connect ecosystem.
| Plane | Product | Owns |
|---|---|---|
| Task | AgentConnect | tasks, artifacts, decisions, reviews, handoffs, worker routing |
| Memory | BrainConnect | the human-gated trusted memory ledger |
| Tool | ToolConnect | which tools exist, who may call them, what happened |
| Compute | ComputeConnect | what compute exists, what it can do, is it healthy, will this fit, where should this run |
AgentConnect and BrainConnect are both consumers of compute. AgentConnect already ships a
LocalComputeProvider contract and an HTTP client for it; BrainConnect's librarian already talks
to an OpenAI-compatible local endpoint. ComputeConnect is the thing that should be on the other
end of both — and one of those contracts is already written, so ComputeConnect conforms to it
rather than inventing a new one.
ComputeConnect can be used on its own. It depends on no sibling project. AgentConnect uses it for orchestration and BrainConnect's librarian routes inference through it (ComputeConnect proxies to an external engine — it never performs or owns the inference itself; see docs/COMPUTE_PLANE.md). Any other application may use it directly:
- Applications and AgentConnect drive the control-plane API (
LocalComputeProvider) — placement, admission, health, cancellation. - BrainConnect and direct applications use the OpenAI-compatible inference API — the same dialect every engine already speaks.
Both surfaces reach one execution backend. The two-API split is specified in ARCHITECTURE.md §5; the stable interface itself lives in docs/CONTRACT.md. No sibling product is required to run or build against ComputeConnect.
A minimal runtime exists (v0.1.0). Two real compute nodes now exist on this host — a 35B MoE engine and a 4B dense engine, differing in family, size, and context window — and single-host heterogeneous placement across them is PROVEN as of 2026-07-27 (real preference-driven selection, capacity-forced placement, and real generation from both engines; see docs/validation/heterogeneity-2026-07-27.md). The simulated cloud provider still exists for the privacy default-deny path only and is never cited as heterogeneity evidence. What remains open is cross-machine heterogeneity — a node of a different physical machine, starting with the GPU-class box on this network — stated plainly in STATUS.md.
| Deliverable | State |
|---|---|
| Product boundaries | Drafted — ARCHITECTURE.md |
| Contracts | Locked — CONTRACT.md: two APIs, five binding invariants; CA-1 and CA-3 implemented, CA-2 proposed |
| AgentConnect contract | Conformed to and tested with AgentConnect's shipped client, including against the real local engine |
| BrainConnect contract | Drafted — a compute consumer on the inference API, not a peer scheduler; nothing rewired yet |
| ToolConnect contract | Provisional — validated runtime but no compute-facing surface yet |
| Code | computeconnect 0.1.0: both API layers, structural privacy, streaming + cancellation, real single-host two-engine heterogeneous placement, 154 tests (143 offline + 11 real-engine) |
| Decisions | D1–D6 all ratified — implementation status per decision in STATUS.md |
| License | Apache-2.0 |
A large fraction of ComputeConnect's originally-stated scope is already owned by mature, actively-maintained, permissively-licensed projects. LiteLLM covers the request plane. Ray and Kubernetes cover in-cluster placement. llama.cpp and Ollama already manage their own model lifecycles. LocalAI is a close analog of the whole idea on a single box.
ARCHITECTURE.md confronts each of these by name and narrows the charter accordingly. The defensible slice is real, but it is much smaller than the initial scope suggested, and the roadmap reflects the smaller slice. A ComputeConnect that drifts into request-level routing collapses into a worse LiteLLM.
The charter is held honest by a ratified design-validation rule (D2): if the demonstrated use case remains a single local host with no heterogeneous placement problem, prefer maintained single-node systems — LocalAI, llama-swap, Ollama, LiteLLM — over building ComputeConnect.
- docs/COMPUTE_PLANE.md — the Compute plane's ecosystem responsibilities: owned/rented/external/marketplace provider distinctions, why Connect never owns or resells the compute, and where cost and budgets live
- docs/ARCHITECTURE.md — boundaries, objects, prior art, the two APIs, the privacy invariant, integration contracts
- docs/CONTRACT.md — the stable interface surface and its future amendments
- docs/ROADMAP.md — phases and the gate for each
- docs/STATUS.md — what is true today, and the ratified decisions
- docs/ORGANIZATION_AWARE_SETUP.md — the Compute plane's part in Connect's organization-aware onboarding: shared-vs-personal compute, residency placement, ownership-vs-authorized-use (design direction)
- LICENSE — Apache-2.0