Ask questions about any codebase. Get answers with exact file and line citations.
AST-boundary symbol chunking + call-path graph tracing. Every answer comes with exact file:line citations — no hallucinated references.
For non-technical readers: Imagine being able to ask "where does this app check if you're logged in?" and getting back a precise, correct answer with the exact file and line number — not a guess, not a paragraph of possibilities. CodeOracle reads a codebase the way a senior engineer does: by understanding the structure of the code (functions, classes, dependencies) rather than just searching for keywords.
Most code Q&A tools chunk source files by token count the same way you'd chunk a PDF. This means a function can be split across two chunks, its docstring lands in a different chunk than the method it documents, and retrieval returns closest-sounding text rather than the actual symbol. CodeOracle chunks at AST boundaries — each indexed unit is a complete symbol: decorator chain, signature, docstring, and full body — parent class context included. Nothing is ever split mid-definition.
📂 Repository Source (.py files)
│
▼
🌳 AST Symbol Chunker
├── Top-level functions and classes
├── Methods with parent class context injected
├── Decorator chain and full signature preserved
└── Cross-file import path extraction
│
▼
🗂️ In-Memory Symbol Index
├── Symbol signature + docstring + body
├── Parent class lineage
└── Import dependency graph (caller ↔ callee edges)
│
▼
🔍 Hybrid Search Engine
├── Dense semantic similarity over embedding space
└── BM25 over tokenized identifiers
(camelCase & snake_case split into subwords)
│
▼
💬 Cited Q&A Generator
Synthesizes answers with exact file:line citations
Identifier Tokenizer — BM25 over raw identifier strings fails badly. Searching for "token limit" won't match compute_token_limit_for_model. CodeOracle splits identifiers into subwords before indexing: parseJwtToken → parse, jwt, token. This enables accurate keyword-style matching without requiring exact naming conventions.
Call-Path Graph Resolution — The import graph built at index time enables transitive dependency tracing. When you ask "what calls this function?" or "what does this module depend on?", CodeOracle traverses caller/callee edges to surfaces answers two levels deep, not just direct references.
Hybrid Retrieval (RRF Fusion) — Dense embedding similarity handles semantic queries ("where is authentication handled?") while BM25 handles exact identifier queries ("show me uses of UserContext"). Reciprocal Rank Fusion merges both ranked lists before passing candidates to the generator.
| Approach | Standard RAG | CodeOracle |
|---|---|---|
| Chunking | Token count windows | AST symbol boundaries |
| Identifier search | Raw string BM25 | Subword-tokenized BM25 |
| Answer evidence | Similar text paragraphs | Exact symbols with file:line |
| Dependency tracing | None | Caller/callee graph traversal |
Query: "Where is JWT authentication verified?"
Answer: JWT authentication is verified in auth/middleware.py at line 47,
inside JWTAuthMiddleware.verify_token(). The method decodes the token using
PyJWT, validates the exp and aud claims, and raises AuthenticationError on failure.
Citations:
📌 auth/middleware.py:47 — JWTAuthMiddleware.verify_token()
📌 auth/schemas.py:12 — TokenPayload dataclass (aud, sub, exp fields)
📌 api/routes.py:83 — @require_auth decorator on protected endpoints
git clone https://github.com/nathaniel-gordon/codeoracle
cd codeoracle
pip install -e .# Query the bundled demo repository
python -m cbe --repo output/demo_repo --query "Where is JWT authentication verified?"
# Interactive REPL session
python -m cbe --interactivepytest tests/ -vcodeoracle/
├── cbe/
│ ├── ast_chunker.py # AST symbol extraction & parent class context
│ ├── identifier.py # camelCase / snake_case subword tokenizer
│ ├── index.py # In-memory symbol index + hybrid search (RRF)
│ ├── call_graph.py # Import dependency graph builder & traversal
│ ├── qa.py # Cited Q&A synthesizer
│ └── __main__.py
└── tests/
Built by Nathaniel Gordon