Bearingkit

A clear workflow for AI coding agents

Bearingkit is an open-source workflow for AI coding agents: one protocol and 17 skills. It works out which step a request belongs to, stops to ask before touching migrations, auth, payments or production, and shows test output before it says a job is done. It runs on Claude Code and Antigravity, and takes requests in English or Vietnamese.

MIT licensed · Pre-release · Developed in the open. Routing has been measured; overall quality and cost are still being tested. Numbers and limits under Status.

An illustration, not a recorded session.

Why I made it

I used to run several skill packs at once: Superpowers, Anthropic's official plugins, a TDD pack, a review pack. Each was good at something, but together they got in each other's way:

  • Two skills would pick up the same request, each with its own rules.
  • Their instructions overlapped and sometimes contradicted each other, and I was the one sorting it out.
  • All of their descriptions sat in the context before I typed anything.
  • Nothing said, across the board, what the agent could do on its own, what it had to ask about, and what counted as done.
  • Moving a project to Antigravity meant setting everything up again.

So I read the sources properly, with my agents doing much of the reading. There were 23 of them. From 22, I listed 1,295 skills, commands, agents and rule files, and marked each one to adapt, keep only as an idea, or drop. Where two sources disagreed, I kept one rule and wrote down why.

What it does

Picks the step

Before reading any code, the agent decides whether a request is a feature, a bug, a review or just a question, and opens the one skill for that job.

Stops for a decision

Tests, behavior-preserving refactors, docs and app-layer fixes: the agent does them and reports. Schema and migrations, auth, permissions, payments, deleting data, remote or production systems: it writes a proposal and waits. If it isn't sure which side a change is on, it asks.

Proves it before “done”

To call a job done, the agent pastes the project's check output; no test means not done. Every claim carries a file:line or is marked unverified, and every number carries its method.

Around those three

  • The chain: classify → spec → plan → build → test → review → ship → close. A small change within three files, outside the risky areas, goes straight to build with tests.
  • Bugs: root cause before any fix; after three failed attempts the agent stops and brings the evidence back for a decision.
  • Hot paths: auth, payments, uploads, migrations and outside APIs get an independent review before push, ideally by a different model.
  • Handoffs: bk-close writes a handoff checked against git; bk-next starts the next session from it, not from memory.

17 skills: bk-spec · bk-plan · bk-build · bk-test · bk-debug · bk-review · bk-ship · bk-close · bk-next · bk-audit · bk-design · bk-map · bk-research · bk-ops · bk-db · bk-perf · bk-setup. Each is described in the README.

Where it stands

17skills plus one protocol
2accepted agents: Claude Code, Antigravity 2.0
96routing test prompts, 44 in Vietnamese
248automated tests in the repository

Bearingkit is pre-release and has been developed in the open since September 2026. It works, but nothing yet shows it does better than the skill packs it was built from.

What has been measured

What has been measured is routing: whether a request opens the right skill. The core set is 60 prompts across six kinds of request (question, small change, feature, bug, review, ship), measured on 2026-09-16 and 17: precision and recall at least 0.9 on both agents; on Claude Code, recall 0.958 and precision 1.000. That run predates bk-db, bk-ops, bk-map, bk-research and bk-perf; prompts for the newer skills are in the 96-prompt set but have not been re-run. Routing correctly says nothing yet about the code that follows. Prompts from the measured set:

  • Add CSV export to the invoices page.bk-spec
  • The login form shows a 500 after a password reset.bk-debug
  • Is this branch safe to merge?bk-review
  • Ship the invoice export.bk-ship

I also compared skills one at a time with the packs they came from: same task, starting code, model and agent, with small samples. Most showed no clear difference; one planning task with one model favored Bearingkit. Where I recorded cost, Bearingkit used about 1.3 to 3 times the tokens of its sources. Against using no skill at all, it cost about 1.9 to 2.1 times on three tasks with Sonnet, and about 7 times on one review task with an Opus reviewer.

Not shown yet

There is no evidence yet that Bearingkit produces better software or uses fewer tokens than the skills it came from. That is what I'm measuring next: Bearingkit against a hand-assembled stack of packs on the same tasks (proposal), and how often the agent acts without asking on trap tasks such as a migration, a deletion or a production push.

Antigravity IDE loads the skills but hasn't passed the acceptance test. Gemini CLI, Cursor, Codex, Copilot CLI and Factory Droid have manifests only and are untested. How each number was measured: docs/measurements.md and the design files in docs/specs.

Install

Once per machine, then activate it in each project that should use it.

# Once per machine: clone the repository (no npm package yet)
git clone https://github.com/tuyenht/Bearingkit.git $HOME/bearingkit

# Claude Code: add the marketplace and install the plugin
claude plugin marketplace add https://github.com/tuyenht/Bearingkit
claude plugin install bearingkit@bearingkit

# Antigravity only: install the skills store
node $HOME/bearingkit/bin/bearingkit.cjs install --host antigravity

# In each project that should use it
cd path/to/your-project
node $HOME/bearingkit/bin/bearingkit.cjs activate
node $HOME/bearingkit/bin/bearingkit.cjs status

$HOME expands in both bash and PowerShell. Per-agent details and uninstalling: docs/hosts.md.

Built in the open

The code, the MIT license, where each part came from (PROVENANCE.md), the measurements and the known limits are all on GitHub. Bearingkit is an independent project, not a product of Anthropic or Google.

I'm Thanh Tuyen Hoang (@tuyenht), a developer in Hanoi and the founder of Bearingkit. I started it in September 2026 and build it with Claude Code. It grew out of my daily work running a small software business, where AI agents do most of the hands-on work. If you use AI coding agents too and want to try it, I'd like to hear what you think, especially where it gets things wrong.