Skip to content

Repository files navigation

ComputeConnect

A compute-resource control plane. It is the authority on what compute exists, what it is capable of, whether it is healthy, whether a given model will fit on it, and where a workload should run.

As of v0.1.0 there is a minimal runtime: the computeconnect package serves both API layers from one backend. Licensed Apache-2.0 (LICENSE). See docs/STATUS.md — including its honest D2 re-evaluation — before proposing work.

Quickstart

python3 -m venv .venv && .venv/bin/pip install -e .   # or: uv pip install -e .
.venv/bin/computeconnect serve --port 8090 --upstream http://127.0.0.1:8080
  • --upstream is an existing OpenAI-compatible llama.cpp server, consumed read-only — ComputeConnect never starts, stops, loads, or unloads it.
  • Port 8090 by default (on the reference host, 8080 is the engine and 8787 is BrainConnect's).
  • Layer 1 (AgentConnect control plane): GET /health, GET /models, GET /models/loaded, POST /route/estimate, POST /generate (streams; returns X-Run-Id), POST /runs/{run_id}/cancel.
  • Layer 2 (OpenAI-compatible): GET /v1/models, POST /v1/chat/completions.
  • Privacy is structural: no privacy_tier means the most restrictive tier — cloud-class providers are filtered before placement and refusals are structured, never silent downgrades.

Tests: .venv/bin/python -m pytest (installs pytest via pip install -e .[dev]; real-engine tests skip when no llama.cpp is reachable on :8080, and the two-engine subset also skips when no second engine is reachable on :8091 via scripts/second_engine.sh). 154 tests total, 143 of them fully offline; see docs/STATUS.md for the current pass count and the real-engine caveat.


The one-sentence version

Every layer of the local-AI stack today assumes compute already exists and is already running: gateways route requests to endpoints somebody else started, engines load models onto hardware somebody else described, schedulers place pods on nodes somebody else registered. ComputeConnect is the layer underneath that answers what is actually out there, right now, and will this fit — and it is deliberately not an inference engine.

What it does not do

ComputeConnect does not perform inference. It never loads a tensor, never picks a quantization, never implements an attention kernel. It knows that llama.cpp on the ARM box can run a 30B MoE, and it knows how to ask it to; it does not know how.

It also does not own tasks, memory, tools, workflow engines, inference engines, or secrets managers. Those belong elsewhere, and the boundaries are drawn in docs/ARCHITECTURE.md.

Where it sits

ComputeConnect is the compute plane of the Connect ecosystem.

Plane Product Owns
Task AgentConnect tasks, artifacts, decisions, reviews, handoffs, worker routing
Memory BrainConnect the human-gated trusted memory ledger
Tool ToolConnect which tools exist, who may call them, what happened
Compute ComputeConnect what compute exists, what it can do, is it healthy, will this fit, where should this run

AgentConnect and BrainConnect are both consumers of compute. AgentConnect already ships a LocalComputeProvider contract and an HTTP client for it; BrainConnect's librarian already talks to an OpenAI-compatible local endpoint. ComputeConnect is the thing that should be on the other end of both — and one of those contracts is already written, so ComputeConnect conforms to it rather than inventing a new one.

Standalone by design

ComputeConnect can be used on its own. It depends on no sibling project. AgentConnect uses it for orchestration and BrainConnect's librarian routes inference through it (ComputeConnect proxies to an external engine — it never performs or owns the inference itself; see docs/COMPUTE_PLANE.md). Any other application may use it directly:

  • Applications and AgentConnect drive the control-plane API (LocalComputeProvider) — placement, admission, health, cancellation.
  • BrainConnect and direct applications use the OpenAI-compatible inference API — the same dialect every engine already speaks.

Both surfaces reach one execution backend. The two-API split is specified in ARCHITECTURE.md §5; the stable interface itself lives in docs/CONTRACT.md. No sibling product is required to run or build against ComputeConnect.

Status at a glance

A minimal runtime exists (v0.1.0). Two real compute nodes now exist on this host — a 35B MoE engine and a 4B dense engine, differing in family, size, and context window — and single-host heterogeneous placement across them is PROVEN as of 2026-07-27 (real preference-driven selection, capacity-forced placement, and real generation from both engines; see docs/validation/heterogeneity-2026-07-27.md). The simulated cloud provider still exists for the privacy default-deny path only and is never cited as heterogeneity evidence. What remains open is cross-machine heterogeneity — a node of a different physical machine, starting with the GPU-class box on this network — stated plainly in STATUS.md.

Deliverable State
Product boundaries Drafted — ARCHITECTURE.md
Contracts Locked — CONTRACT.md: two APIs, five binding invariants; CA-1 and CA-3 implemented, CA-2 proposed
AgentConnect contract Conformed to and tested with AgentConnect's shipped client, including against the real local engine
BrainConnect contract Drafted — a compute consumer on the inference API, not a peer scheduler; nothing rewired yet
ToolConnect contract Provisional — validated runtime but no compute-facing surface yet
Code computeconnect 0.1.0: both API layers, structural privacy, streaming + cancellation, real single-host two-engine heterogeneous placement, 154 tests (143 offline + 11 real-engine)
Decisions D1–D6 all ratified — implementation status per decision in STATUS.md
License Apache-2.0

The honest risk

A large fraction of ComputeConnect's originally-stated scope is already owned by mature, actively-maintained, permissively-licensed projects. LiteLLM covers the request plane. Ray and Kubernetes cover in-cluster placement. llama.cpp and Ollama already manage their own model lifecycles. LocalAI is a close analog of the whole idea on a single box.

ARCHITECTURE.md confronts each of these by name and narrows the charter accordingly. The defensible slice is real, but it is much smaller than the initial scope suggested, and the roadmap reflects the smaller slice. A ComputeConnect that drifts into request-level routing collapses into a worse LiteLLM.

The charter is held honest by a ratified design-validation rule (D2): if the demonstrated use case remains a single local host with no heterogeneous placement problem, prefer maintained single-node systems — LocalAI, llama-swap, Ollama, LiteLLM — over building ComputeConnect.

Documents

  • docs/COMPUTE_PLANE.md — the Compute plane's ecosystem responsibilities: owned/rented/external/marketplace provider distinctions, why Connect never owns or resells the compute, and where cost and budgets live
  • docs/ARCHITECTURE.md — boundaries, objects, prior art, the two APIs, the privacy invariant, integration contracts
  • docs/CONTRACT.md — the stable interface surface and its future amendments
  • docs/ROADMAP.md — phases and the gate for each
  • docs/STATUS.md — what is true today, and the ratified decisions
  • docs/ORGANIZATION_AWARE_SETUP.md — the Compute plane's part in Connect's organization-aware onboarding: shared-vs-personal compute, residency placement, ownership-vs-authorized-use (design direction)
  • LICENSE — Apache-2.0

About

Modular AI compute orchestration for local and cloud execution. Manage runtimes, providers, hardware, and workload placement through pluggable adapters.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages