Developer guide¶
mokata is a pure-Python package under src/mokata/, with one capability model. Supported
Python: 3.10–3.13.
Dependencies. The one required runtime dependency is the MCP SDK (mcp>=1.2,<2) — it
ships by default so the bundled mokata-mcp server works straight out of pip install mokata.
The upper bound is deliberate: mcp 2.0.0 removed mcp.server.fastmcp, which the mokata MCP
server is built on, so 2.x needs a port rather than a version bump. If your environment already
has mcp 2.x, mokata-mcp says so and tells you to run pip install 'mcp<2'.
Everything else is optional and degraded over when absent, never fatal:
| Extra | Pulls | Absent ⇒ |
|---|---|---|
schema |
jsonschema>=4.0 |
the built-in structural validator still validates the manifest |
postgres |
psycopg>=3.1 |
memory degrades to the SQLite floor |
embeddings |
model2vec>=0.3, numpy>=1.24 |
the semantic tier falls to the zero-dep hashing floor (token-hash overlap, not meaning) — mokata doctor reports which tier is actually ranking recall |
mcp |
(no-op alias of the default dep — kept so mokata[mcp] still resolves) |
— |
Every optional import is lazy, so the core, the CLI and every default profile run with all of them missing.
Architecture by package (Parts A–L)¶
The spine is the conductor; every other layer plugs into it through the manifest, the
capability router, and the unified Surface.
Part A — Spine¶
manifest.py— load/validate the stack manifest (Manifest), accessors for layers, capabilities, tools, settings;layer_enabled,tool_enabled,capability_enabled.schema.py— structural validator (authoritative, dependency-free) + an optionaljsonschemapass that degrades on any failure.detect.py—Detector: is a tool present? (command/python_module/path/always), with overrides + caching. Absence is a value, never an error.router.py—Router.resolve(need)walks a capability's declared fallback order and returns the first present provider, recording the attempted chain (Resolution).config.py—Surface: the single governed read surface over.mokata/(manifest + constitution + router + state store).bootstrap.py— the SessionStart briefing, capped at a 2,000-token budget.init.py/profiles.py/cli.py—mokata init, the tool catalog + profiles, the CLI.adapters/— A6/H4–H6:AdapterContract+negotiate(coverage/gaps),MCPRegistry(discovery),overlapping_capabilities/resolve_conflict(precedence).
Part B — Knowledge (knowledge/)¶
query.py (typed QueryResult/Reference, 5 query kinds), grep_backend.py (the
lexical floor), ast_backend.py (the embedded stdlib-AST floor — non-degraded structural
queries on Python, a floor above grep), graph_backend.py (the adopted code-review-graph
adapter via an injected client), layer.py (KnowledgeLayer — backend chosen through the router, story bridge),
index.py (incremental fingerprint index + staleness surfacing), anchors.py (@lat
drift anchors + lat_check).
Part C — Memory (memory/)¶
item.py (MemoryItem + the three types), backends.py (SQLiteBackend default,
PostgresBackend), store.py (the logic: gated writes, toggles,
instrumentation, consolidation), healing.py (surfacing detection), episodic.py
(searchable turns, lexical fallback), consolidation.py (proposal-only).
Retrieval is tiered and honest about it: tiered.py (the ranking pipeline), embed.py +
vector.py (the opt-in semantic tier — model2vec via mokata[embeddings], else the zero-dep
hashing floor), reembed.py (the gated re-embed that resolves a stamp mismatch), and
tier_report.py (the retrieval stack lines mokata doctor prints — which engines are actually
ranking a recall). The lexical tier runs in the database where the backend supports it (SQLite
FTS5/bm25, Postgres tsvector/ts_rank) and degrades to a Python Jaccard floor otherwise.
share.py is the .mokata/backups/ export/import surface; migrate.py ports the live store
between backends; review.py is the Draft→Published proposal workflow.
Part D — Engine (engine/)¶
spec.py, acmapper.py (AC → test traceability), completeness.py (the blocking gate),
premortem.py (risk probes), phases.py (analysis/strawman + run_pipeline),
compliance.py (spec-compliance review), preview.py (zero-side-effect dry-run).
Part E — Execution (execmode/, modes/)¶
selector.py (per-run mode choice), tasks.py, orchestrator.py (isolation, fan-out,
handback cap, degrade), review.py (two-stage), routing.py (cheapest-capable model +
escalation); modes/bug.py, modes/debug.py, modes/optimize.py.
Parts F/G/I — Governance (govern/)¶
tokens.py, retrieval.py, compaction.py, compress.py, budget.py, cache.py (F);
rules.py, karpathy.py, learning.py, authoring.py, hooks.py, enforce.py (G);
secrets.py, gate.py (WriteGate + trust enforcement), ledger.py (hash-chained),
trifecta.py, deviation.py, outbound.py, revert.py, resume.py, tdd.py (I);
trust.py, doctor.py, lifecycle.py (K).
The seatbelt — enforcement outside our own tools¶
The gates above all fire inside mokata's tools; these modules are what stop the model simply
reaching past them. Read gate_hook.py's module docstring first — it is the design record.
approval.py— the human-minted approval:propose/redeem, content-hashed and session-scoped proposal ids.approve=trueon a tool call is inert by construction; onlymokata approve <id>(a human, out-of-band) mints one, and it licenses exactly one commit.gate_hook.py— the decision for the four run-state gates (approach-approval,spec-persisted,no-code-without-failing-test,spec-scope) on a nativeWrite/Edit. Pure, total, never raises; every uncertainty (no run, ambiguous run, unreadable state, undeclared scope) resolves to ALLOW.hook_cli.py— the I/O for all three shipped hooks (session-start,secret-guard,gate-guard), launched via themokata-hookconsole entry point. Blocks with exit code 2.spec_scope.py— a spec's authorized surface + its deferred items (paths and literal markers), andclassify(), the pure verdict the scope gate reads. Plus the amend record.tdd_state.py/session_state.py/session.py— the persisted RED/GREEN record, the run-scoped (__<run_id>) state keys, and the minted per-process session identity that makes two windows on one repo distinguishable.degrade.py— the degrade registry: a capability that falls back to a floor is remembered, sodoctorcan say so instead of letting a silent fallback pass for the real thing.
Parts J/K/L — Distribution & composability¶
harness.py (thin cross-harness boundary), harness_setup.py (mokata setup <harness> —
commands + MCP + the three hooks), share.py (export/import stacks), compose.py (chaining +
suggestions), playbook.py (the end-to-end integration runner), packaging.py
(plugin/marketplace validators), team*.py (the opt-in shared Postgres store).
The MCP server (mcp/)¶
registry.py is the single source of truth for the tool set — 61 tools: 40 read, 20 write, 1
approve — and server.py builds the mokata-mcp server from it. Every write tool is
propose-only; tools_approve.py is the default-off, opt-in in-chat approval
(settings.approvals.in_chat). The robustness layer is shared by every tool rather than
re-implemented per tool: tool_annotations.py (read-only / destructive hints),
response_format.py, pagination.py, and validation.py (input validation at the boundary).
../awaiting.py supplies the "what is waiting on YOU" head that the statusline, the notification,
and mokata doctor all render from one place.
Dev setup¶
git clone https://github.com/JasGujral/mokata-oss && cd mokata-oss
pip install -e ".[schema]" # editable install + jsonschema; the MCP SDK is a default dep
(End users never clone: pip install mokata → mokata setup claude.)
Running the tests (BOTH jsonschema states)¶
A hard invariant: the suite must pass with jsonschema absent and present.
# absent
pip uninstall -y jsonschema
python -m unittest discover -s tests -t tests
# present
pip install "jsonschema>=4.0"
python -m unittest discover -s tests -t tests
CI runs both jsonschema states on ubuntu-latest and windows-latest, on Python 3.12,
plus a mokata playbook smoke run — and the declared floor on three further legs (ubuntu ×
present, ubuntu × absent, windows × present). The matrix is deliberately light; the classifiers
cover 3.10–3.13. See platform support.
Tests are written RED-before-GREEN.
On the declared floor¶
python -m unittest above runs on whatever interpreter is on your PATH, which on many machines is
older than the floor the package promises. Do not guess — provision it:
scripts/floor-python.sh # build build/floor-venv at the declared floor
scripts/floor-python.sh --jsonschema absent # the same, with jsonschema removed
scripts/floor-python.sh --dry-run # both routes, and whether this machine has either
scripts/floor-python.sh --check # what version is actually in there?
scripts/floor-python.sh --exec -m unittest discover -s tests -t tests
The floor is read from pyproject.toml's requires-python on every run, never hard-coded, so the
command keeps working when the floor moves. --check distinguishes at the floor, below it,
above it and not provisioned with four separate exit statuses: an interpreter newer than the
floor is refused too, because a run on it is not a floor run.
Provisioning needs one of two routes: uv, which fetches the
interpreter itself, or a python<floor> already on your PATH. --dry-run prints the command each
route would run and marks each available or absent; if you have neither it exits 2 and says
so, rather than printing an empty plan and succeeding.
Contributing¶
See CONTRIBUTING.md
for the full flow. The non-negotiables: TDD (RED-before-GREEN), clean-room (no import
of or text from any other framework), human-gate every durable write, local-first, and
Apache-2.0 / MoStack with no vendor-prefixed names. To add a skill/command, register a
Skill in skills.py and regenerate its template; to add a tool, declare an
AdapterContract and wire it through the router.