The boost roadmap
Every finding from the autonomous quality loop lands here — what shipped, what's mid-flight, and what's queued — scored by complexity, impact, and a little bit of wow.
Shipped
// merged to main & publishedGitHub Pages deploy is broken every push
Mutation hardening — core/store.py
Mutation hardening — core/gitutil.py
Untrack generated build noise
browse crashes when you pick a rule or workflow
Consolidate skill-staleness / drift logic into core
Write-up · consolidate-skill-staleness-drift-logic-into-cor.md
Mutation hardening — core/frontmatter.py
Extension-free tests — core/dense.py
Crash-recorder error paths — core/logs.py
Extract MCP + HTTP servers out of configuration.py
Write-up · extract-mcp-http-servers-out-of-configuration-py.md
Autonomous ship-workflow & isolated worktree
Write-up · autonomous-ship-workflow-and-isolated-worktree.md
Fix self-update version detection (dead branch)
End-of-options -- guard on git commands
Crash-correlation breadcrumbs in the invocation log
The copyleft protected nothing and cost the one thing boost needs
bmad on knew about a second host and wrote to one anyway
A Gemini user got the heuristic fallback from every AI command
Reuse helpers; kill minor dead work
All 80 boost commands audited against a disposable HOME — four defects fixed
fix_hint's no-key guard has been unreachable since the day it was written; a missing API key now prescribes the full re-embed the guard exists to prevent
The eval corpus's size and its concentration ceiling are counted with len(scan_dir) — the measure measure_registry.py exists to say is wrong — so 44.7% of the gate's corpus is vendored …
Write-up · eval-corpus-counted-by-the-method-the-repo-disavows.md
The 2026-09-01 pin refresh moved the corpus and re-baselined it, but nothing re-derives the floors — CLAUDE.md's "~10% under measured" is now 7.2%–17.3%
Write-up · eval-floor-calibration-stale-after-pin-refresh.md
The exemplar mechanism was applied to the ungated set only: golden.jsonl is 0/91 pinned, and 10 of its 44 hit@1 credits are on names the metric cannot adjudicate
The exemplar migration never reached the query set the required gate runs: 0 of 91 rows, and 27 of them have an ambiguous target on the gate's own corpus
Research: how boost install could tell that a skill needs something else installed first
A locally built vector store displaced by a new API key is told to re-embed through the paid provider, where unsetting the key is free
Next up
// triaged findings, starting soonReconcile the theme drift
Planned
// on the listFinish mutation hardening across core/
Bring commands/ under mutation testing
The ~8,100-line command layer has zero mutation coverage: mutmut is
scoped to core/ only. A blocking 80% floor was attempted and the
baseline has now actually been measured — it does not clear the bar, and
two of the three constraints recorded here earlier were wrong. What a sample run
of commands/taps.py (the best-covered module in the package,
95.8% lines) against tests/unit/ + tests/functional/
reports through this repo's own gate:
570/793 killed — 71.9%, under the 80% floor, with all 223
survivors spread across every one of the module's 7 functions rather than
concentrated in one testable gap. Mutant density is 4.2 per
statement in taps.py but 2.46 across
core/ (9,909 mutants over 4,033 statements), so
commands/'s 5,314 statements imply somewhere between
13,000 and 22,300 mutants; at the measured
1.42 mutations/sec — roughly 6.5× slower per mutant than the
core/ job, which matches the functional-vs-unit suite cost — that is
2.5 to 4.5 hours on one runner, against an ~18-minute job today.
The job sets no timeout-minutes, so it inherits GitHub's 360-minute
cap rather than failing fast.
Two blockers sit underneath that number. Selecting only tests/unit/,
the way core/ does, leaves commands/ at
17.9% line coverage versus 91.5% with functional included — and
"no tests" mutants count against the score, so that route floors out near 18%.
Selecting tests/functional/ instead crashes the run: mutmut's
record_trampoline_hit calls p.resolve(strict=True) on the
relative source path, so any test that chdirs into a temp
project dies with FileNotFoundError: <tmp>/boost_cli — confirmed
on test_verify_sees_a_project_skill. Corrections to the earlier note:
pytest_add_cli_args_test_selection is not single-valued, it
is a list that configparser splits on newlines (a space-separated line is what
errors), and a sibling pytest_add_cli_args takes extra pytest flags.
One prerequisite is already fixed: --no-mcp leaked through
os.environ and made the suite order-dependent, which mutmut exposes
because it runs whichever test subset covers each mutant.
Declined 2026-07-29. Not "hard" — blocked, and the block is
upstream. mutmut 3.6.0 is still the latest release, and
src/mutmut/__main__.py:120 still reads
source_paths = [p.resolve(strict=True) for p in Config.get().source_paths]
— unconditionally, before the max_stack_depth guard — on paths that
configuration.py:102 builds as plain relative Paths. Every
chdir-ing functional test therefore kills the run, and the functional
suite is precisely what takes commands/ from 17.9% to 91.5% line
coverage. So the two candidate configurations are "floors out near 18%" and "crashes";
there is no third. Even granting a fix, the measured numbers already refuse the
proposal on their own terms: 71.9% against an 80% blocking floor, with
survivors spread evenly rather than pooled in one testable gap, at
2.5–4.5 hours per run against an ~18-minute job. A gate that is 8 points red
on the day it lands does not gate anything; it just makes main red.
The lead, if anyone reopens this: the crash is a relative-path bug, and
source_paths is not required to be relative — an absolute path survives
resolve(strict=True) from any working directory. Whether the rest of
mutmut's mutants/ copy machinery tolerates one is untested. That, plus a
non-blocking scheduled job that rotates one module per week, is the shape worth
trying; a blocking 80% floor is not.
What is not lost by declining. The architecture already puts behaviour in
core/ — which is mutation-gated at 80% — and keeps
commands/ as thin CLI glue. The uncovered layer is the one deliberately
designed to hold the least logic, and it still carries 91.5% line coverage from the
functional suite.
Visual regression pass on the guide
Refresh the marketing surface
a Tier 3 eval for tool-call behaviour, floored in both directions
boost's required gate floors four retrieval metrics — recall@k ≥ 0.78, hit@1 ≥ 0.40,
MRR ≥ 0.52, nDCG@k ≥ 0.58 — over a 91-query golden set and a 10,152-entry corpus, every
row pinned to a commit SHA. All of it measures what boost returns once it is asked. Nothing
measures whether an agent asks. **The call itself is unmeasured**, and it is the step everything
downstream depends on.
The miss that exposed it. A Gemini CLI session was asked to "create a new, simplified app
demonstrating RAG implementation in Python3 using langGraph, langChain, and langSmith" — a new
project, an architecture decision and a dependency choice, which is three of the triggers
boost_search's description names explicitly. It activated two already-installed
skills, built the app, and never called boost. Asked why, it paraphrased boost's own lock-in
trigger list back verbatim, so the text was read and was not persuasive. A gate that floors
recall@k at 0.78 reported nothing, because retrieval was never invoked.
Every claim in the MCP surface is argued, not measured. The triggers, the 10-15s stated
cost, the skip list, the three-kind framing, the "already covered is not already checked" defeater
— each survived a careful review and none has a number behind it. That is a
high-variance lever tuned blind: Tool Preferences in Agentic LLMs are Unreliable
(EMNLP 2025, arxiv 2505.18135) measures description-only edits swinging call rate by
more than 10×. boost currently ships those edits on reasoning alone.
The design constraint that decides whether this is worth building: floor both directions.
A tier that measures call rate alone rewards making boost maximally assertive, which is precisely
the capture the surface is written to avoid — and boost has already learned this exact lesson one
tier down. Flooring recall alone was a hole rather than a simplification: a ranker that
finds the right answer every time and never ranks it first scores recall@10 1.000 with hit@1
0.000, and passed. So the prompt set needs two halves — a should-call set (multi-file work,
a new subsystem, a config or CI job that outlives the session) and a should-not-call set
drawn from the shipped skip list (a question, a one-line edit, a command the user just handed
over) — with a false-call ceiling as binding as the call-rate floor. One number without the other
is an incentive to ship the thing boost refuses to be.
Per host, never averaged. The two registered hosts do not see the same boost text. Claude
Code puts server instructions in the system prompt; Gemini CLI never delivers them in
interactive mode at all — Config.initialize() does not await
mcpInitializationPromise, so getMcpInstructions() returns "",
startChat stamps the context entry once with a stable id, and the later
refreshMcpContext() re-renders Tier 1 only. A single averaged score would hide a host
where 1,786 characters of guidance are simply absent, and would credit or blame wording for a
delivery failure.
Shape. An opt-in make eval-tools beside eval-ai /
eval-rec / eval-explain — real hosts and real LLM calls, so it is
non-deterministic and key-gated and must not join the required check gate;
same degrade-cleanly contract as the other Tier 2 evals. Because the outcome is stochastic, report
N runs per prompt with an interval rather than a single pass/fail, the way
golden-set-statistical-power established for retrieval — a one-shot replay cannot tell
a wording regression from a sampling wobble.
Where this stands (2026-08-31), and a correction. The Claude Code arm shipped in #616:
scripts/eval_tools.py, a 16-prompt set halved into should-call and should-NOT-call,
Wilson intervals over N runs, and a verdict that floors call rate and ceilings false calls.
Its probe was broken, and the finding this card recorded was an artifact of it. An earlier
revision of this card reported “3/3 false calls — boost's tools fired on What is the
difference between a Python list and a tuple?”. They did not fire. called_boost()
substring-scanned the raw event stream, and claude -p --output-format stream-json --verbose
opens with a system/init event enumerating every tool available to the
session — which on any machine where boost is registered contains
mcp__boost__boost_search and the other three CONSULT names. So the check returned
true on every run, including runs with no tool call at all. Measured directly:
Say OK and nothing else. produced zero tool_use blocks and scored as a
boost consult.
Two consequences shipped with it. make eval-tools could never pass — eight
no-call rows × three runs is 24 forced trues, so the false-call rate's lower bound sat at
1.00 against a 0.20 ceiling, red on every machine forever. And the tier built to retire
unfalsifiable claims had produced one. The lesson is the tier's own: the existing tests
passed because they fed hand-written one-line fragments with no init event — a
fixture the author invented could not catch the author's wrong model of the input. The probe now
parses the NDJSON and counts only tool_use blocks inside assistant
events, and the regression test drives a captured real stream.
The second host arm is still unwritten, and should stay that way until the fixed probe is
re-run. Building arm two on a probe that cannot tell an offer from a call would produce two
hosts scoring an identical, meaningless 1.00. When it is built, the candidate is Gemini CLI
proper, not Antigravity CLI: the delivery claim below is about Gemini's Node bundle, and
agy is a third mode again — it receives boost's instructions and
writes them to ~/.gemini/antigravity-cli/mcp/boost/instructions.md, pointing the agent
at the file rather than inlining it. Substituting it would measure a different mechanism than the
one this card argues about.
Cost, now measured. Two trivial runs on a real host reported $0.657 and
$0.682 of total_cost_usd, so 16 prompts × 3 runs is roughly
$30–50 per host per invocation on a machine with a crowded tool surface.
--strict-mcp-config with a boost-only config cuts that sharply and controls the
surface confound in the same move.
2026-08-31: --strict-mcp-config shipped. eval_tools.py now takes
a --strict-mcp-config flag: it writes a boost-only mcpServers config
(the same <launcher> mcp --stdio invocation and fork-safety env
core.mcphost.register_argv uses for a real registration — confirmed against an actual
claude mcp add-json write, not guessed at the schema) to a temp file and passes
--strict-mcp-config --mcp-config <path> to every claude -p call,
cleaning the file up afterward. This session's sandbox had no network path to PyPI, so the pinned
toolchain (pytest, ruff, mypy, …) could not be installed and
make check could not be run here; the change was verified by hand instead — direct
python3.12 import of the module, the new unit tests executed by eye against the
interpreter, py_compile, a manual line-length check against ruff's 88-column default,
and an end-to-end dry run with subprocess.run mocked that confirms the flags land on
the argv and the temp file is created and removed. CI runs the real gate on the PR.
Still unwritten: the second host arm (Gemini CLI). No gemini CLI was reachable
in this sandbox to capture a real stream from, and building that arm on an invented model of
Gemini's non-interactive output format is the exact mistake this card's own probe fix (2026-08-30)
already paid for once — "a fixture the author invented cannot catch the author's wrong model of the
input." That arm stays a placeholder until it can be built against a captured real stream, on a
machine with the gemini CLI installed.
What it unlocks. The first honest answer to "did that description edit help", a baseline the
next surface change can regress against, and a way to retire claims that survive only because
nobody can check them.
unpin the [eval] langchain stack when ragas ships its fix
The [eval] extra pins langchain-core<0.4,
langchain-community<0.4 and langchain-openai<1 because ragas hard-imports
ChatVertexAI from a langchain_community chat-models path that 0.4.x
deleted. The LangChain integration card originally made this unpin its phase 0 and was
corrected in place: ragas 0.4.3 still carries the import (measured 2026-08-04 — declared
bounds are open, but import ragas crashes beside langchain 1.x), while upstream main
already has the removal merged. So the unpin is one release of someone else's package away.
What to do when it lands. Check pip index versions ragas (or the PyPI JSON) for
a release after 0.4.3; verify in a throwaway venv that import ragas succeeds beside
langchain>=1; then move [eval] to that floor, delete the three langchain
pins, and adapt scripts/eval_explain.py if the 0.4 scoring API moved (its
evaluate/to_pandas surface is what
test_eval_faithfulness.py stubs in the unit suite). The eval-explain workflow is
the live proof — it must stay green with real keys.
Re-checked 2026-08-30. pip index versions ragas still reports
0.4.3 as the newest release, so nothing has changed and this card is still not
claimable. Recorded here rather than left implicit: a card that says "check before starting"
gives a reader no way to tell a check that came back negative from a check nobody ran.
Why it stays its own card. The shipped integration card documents the block but will not be
re-read; an unpin nobody remembers is how a workaround pin outlives its reason by years. This card
is the reminder, and it is deliberately not claimable until the upstream release exists.
give boost-langchain a release path to PyPI
Declined, deliberately. Both missing pieces below were owner-only or upstream-blocked, and
the research they prompted dissolved the premise: a second PyPI project bought a separate release
cadence nobody needed (boost releases more often than langchain), while the ecosystem evidence —
langchain-community sunset, non-langchain-* names in LangChain's own integrations
listing, in-host precedents from ragatouille to mlflow — showed the
standalone distribution was never required. The integration now ships inside the
boost-skill-cli wheel behind a [langchain] extra instead; see
langchain-in-the-wheel. The original card follows for the record.
The boost-langchain distribution shipped under integrations/langchain/
with its whole point being a separate release cadence from boost-skill-cli —
langchain majors move faster than boost does, and the conformance workflow already builds the
sdist/wheel and runs twine check on every touching PR. What does not exist is any way
for those artifacts to reach PyPI: the name 404s there, and nothing publishes on any trigger.
Two pieces, one of which only the repo owner can do. First, create the PyPI project and
configure a Trusted Publisher for it — pending-publisher registration works before the first
upload, and the filename-matching rule that pinned boost's own workflow name applies here too.
Second, a publish workflow with a deliberate trigger: not boost's every-merge cadence
(publish.yml releases boost-skill-cli on every push to main, which is
exactly the coupling the separate distribution exists to avoid) — a tag like
boost-langchain-v0.1.0 or a manual dispatch that bumps the static version, builds from
integrations/langchain/, and publishes with the OIDC token. Remember the repo's own
lesson: a release:-triggered workflow can never fire here (GITHUB_TOKEN events do not
chain), so trigger on the tag push or dispatch directly.
The floor is already honest. The package requires boost-skill-cli>=1.0.320 —
measured against the actual API it calls, verified by an adversarial install — so the first
published version works against PyPI as it stands today.
publish the keyword index the way vectors are published
Dense vectors are built once in CI and downloaded. The BM25 index is not:
core/rag.py has no export or import function at all, and
shards.yml / scripts/publish_shards.py are dense-only end to end. Every
install rebuilds the same index from the same registries, at the same pinned commits, to produce
the same bytes.
Measured, on a real 458-tap machine. The on-disk index is
rag_index.json 43.7 MB plus rag_postings.sqlite
653.0 MB — 696.7 MB for 18,619,658 postings. Build cost, timed over a
9,306-entry / 69-tap slice: 4.54 s reading bodies and tokenizing, 3.83 s
writing postings, 8.4 s total — about 0.9 ms per entry, so roughly
65 s and ~900 MB extrapolated to the full 71,700-entry catalogue.
Which user actually pays it. Not the default one: boost quickstart taps the
7 starter registries and indexes them in about a second. The cost lands on
boost quickstart --catalog — 463 registries, 2 min 10 s of parallel cloning
and then a minute of indexing on top — and on anyone who taps their way there gradually.
The bug that makes this worth doing is not speed. boost catalog --import
already exists and already looks like the answer: shareable-catalogue-bundle advertises
10.9 MB replacing a 12 GB clone and "59,972 searchable items in 4 seconds". That 4
seconds is fast for a reason the card does not state. rag.read_body degrades
silently to name + description when the item's clone is absent
(rag.py: "Missing files degrade to just the catalog metadata"), and a bundle import
restores catalogues with zero repositories cloned. So the index it builds is not the
full-content index the evals gate floors — it is a frontmatter index wearing the same
file name.
Measured directly over 3,015 real entries, indexing them with and then without their
clones: 3,041,326 tokens versus 182,507. A bundle-only index carries 6.0% of the
searchable text, and nothing in the output says so. That is the same failure shape as an
unpinned eval corpus — a number that still renders confidently while measuring something else.
Why this is easier than the dense shards, not harder. BM25 looks like it needs global
statistics, and it does — but none of them are frozen at build time. _bm25 derives
n = len(docs) and df = len(plist) on every query, so IDF is
computed from whatever corpus is loaded. A per-registry shard therefore merges by offsetting
doc_id, unioning the postings, and recomputing avg_len from per-shard
totals — arithmetic, not re-derivation. And unlike vectors there is no embedding space to match
and no API key to hold, so shards.incompatible() has no analogue here: a published
keyword index is importable by everyone, including the keyless user who cannot use vectors at all.
Shape. rag.export_shard / rag.import_shard mirroring
dense's pair, per-registry assets on the existing shards-latest release,
rows carried in the same manifest.json with the same commit pin and sha256 — the
carry-forward machinery in publish_shards.py manifest --carry-forward applies
unchanged, because a registry whose commit did not move has an index that did not change either.
Three invariants transfer verbatim from the dense side and each is load-bearing: verify before
replacing, refuse a shard whose commit is not the tap's commit, and never treat a missing digest
as a match.
The open question is payload size, and it is large enough to be its own decision — see
shrink-the-published-index. This card should not ship
until that one has an answer, because publishing 697 MB per refresh to save 65 s of CPU
is not obviously the right trade, and at the compressed sizes measured there it clearly is.
Partly landed, and deliberately still inflight — 2026-09-10. What shipped is
the half that needed no size decision: the index now records what it is.
read_body_full returns the text and whether it contains the item's body,
build() reports metadata_only over every document written (reused ones
included, or an incremental build reports zero on the run after a bundle import),
index_completeness() reads the share back off disk, and boost reindex
says it out loud instead of reporting the same confident count for a 6% index. The share is of
tokens, not documents: a bodyless entry still produces a document, so a document share sits
at 1.0 until it drops to 0.0. INDEX_VERSION moved to 9, because the flag is
written only when a body is missing and absence may only be read as "complete" once no older
document can survive.
What did NOT land: rag.export_shard / rag.import_shard, the
per-registry assets, and the manifest.json rows — the publishing pipeline itself.
That half is what the payload-size question governs, and
shrink-the-published-index still has no answer: its claim
is stale, not active — branch loop/shrink-postings-index was last touched
2026-09-02, carries one commit, has no pull request, and is 394 commits behind
main. Someone should un-claim it. One structural finding for whoever takes it: doc
ids are positional (_save does enumerate(docs)), so the card's
"merge by offsetting doc_id" is sound as written — and the shard format should
serialize logical postings (digest → term → tf) rather than the SQLite layout, so the
interning that branch was attempting cannot invalidate a published shard.
shrink the keyword index before publishing it — structure first, then compression
publish-the-keyword-index is worth doing only if the
artifact is small enough to ship weekly. This card is the measurement that decides it, and the
first answer is that compression is the second lever, not the first.
What the format actually stores. _write_postings creates
postings (term TEXT, doc INTEGER, tf INTEGER) and inserts one row per posting, so the
term string is repeated in every row. Measured on the real 458-tap store:
18,619,658 rows over 210,422 distinct terms averaging 6.6 characters. That is
~123 MB of term text to carry 1.4 MB of distinct term text — 88× redundancy,
before the per-row and B-tree overhead that turns it into a 653 MB file (page_size 4096,
167,174 pages, freelist 0, so it is not slack space).
Compression measured on that file, as it stands:
rag_index.json 43.7 MB raw · gzip -6 10.6 MB (4.12×).
rag_postings.sqlite 653.0 MB raw · gzip -6 201.7 MB (3.23×) · zstd -3 169.9 MB (3.84×) · zstd -19 106.7 MB (6.11×).
So even with no format change, zstd -19 puts the whole index near 117 MB — under half
the ~300 MB of dense vectors already published weekly. The trade is already good; the point
of this card is that it can be much better, and that the two levers compose.
Structure first, and it is the bigger win. Interning terms into
terms(id, term) with postings(term_id, doc, tf) removes ~123 MB of
duplicated strings and shrinks the postings_term index from a text key to an
integer one. Beyond that, the classic inverted-index encodings apply directly because doc ids
within a term are ascending: delta-encode them, varint or bitpack the deltas, and store one blob
per term rather than one row per posting. Both shrink the file on disk, not just in
transit, which is the half a compressed download never gives back — the user still ends up with
653 MB resident after import.
What must not regress. read_postings exists precisely so a query touches a
handful of terms instead of materialising the whole map — the change that took cold search from
8-13 s and multiple GB resident to 31-70 ms of scoring. A blob-per-term layout keeps
that property (one row read per query term, decoded on the spot); a scheme that requires decoding
neighbouring terms to find one does not. _bm25 must stay byte-identical, as it did
through the SQLite move, and TestBm25Math is what says so.
Decompression cost is the thing to measure, not assume. zstd -19 is slow to compress and
fast to decompress, which is the right asymmetry for a weekly build feeding many imports — but
"fast" needs a number on the import path before it is a claim, next to the 0.12 s that
importing dense rows costs today. A zstd dictionary trained across shards is the obvious follow-on
for the many-small-registries case, where per-shard compression has little context to work with.
Deliverable. A measured comparison — raw, interned, delta+varint, each × none/gzip/zstd
— on the real store, with import-side decode time beside each. That table is what tells
publish-the-keyword-index what to ship, and it is worth
having even if publishing is declined: the on-disk win applies to every install today.
Progress — PR 688, merged as f003fa03 in train 691. The structural half shipped: _write_postings now
interns terms into their own terms(id, term, df) table, with postings
carrying an integer term_id instead of repeating the term string on every row —
exactly the "structure first" change this card calls the bigger win, and it bumps
INDEX_VERSION so every store picks it up on its next rebuild.
stem_expansions now reads the precomputed df column directly instead of
a GROUP BY COUNT(*) over postings on every prefix lookup. Not done: the
delta/varint doc-id encoding, and the full raw/interned/delta+varint ×
none/gzip/zstd comparison table with import-side decode times on a real multi-hundred-MB store —
this sandbox has no such store to measure against, only a small synthetic one (interning alone cut
a 1.8M-posting/20k-term synthetic store from 63.6 MB to 47.9 MB, directionally consistent
with the real-store estimate above but not a substitute for it). Left as follow-on work before this
card can be called shipped.
A best-effort log handler prints a traceback over every command's output
Near-identical copies survive content-hash dedup and take the whole result page
Content-hash dedup shipped and worked:
rag.dedupe_by_content took duplicate result slots from 4.94 to 0.60 per query
over a 77-tap corpus. That card closed naming one thing still open — near-identical
rather than byte-identical clustering, where core/typosquat.py's confusion machinery
would apply — and buried it under a shipped status where nobody would claim
it. This card is that remainder, with a measurement that makes it look considerably worse than
“refinement”.
Observed on a real 466-tap install with hybrid RRF serving (658,131 chunks): for the query
exa search, every one of the top ten rows is exa-search, and the
descriptions are what give the
shape away — one Japanese (Exa MCPによるウェブ、コード、企業調査), two Chinese
(通过Exa MCP进行神经搜索), five English variants of Neural search via Exa MCP,
plus Use Exa MCP for current web… and AI-powered web search….
All ten are ★ curated. The footer reads
51 matches · ranked by hybrid RRF (BM25 + dense).
Every one of those passed dedup correctly. They are not byte-identical: they are the same
skill in Japanese, in Chinese, and in five English phrasings across different registries. The body
digest differs, so dedupe_by_content keeps them all — which is exactly the
behaviour #366 proved must be preserved, since two entries sharing a name can be
genuinely different rules. The shipped fix is not misbehaving. It simply does not reach this shape.
What the 0.60 residual actually was. The prior card described its leftover as “entries
sharing a name whose bodies genuinely differ, which must stay separate” — true
as stated, and it reads as a rounding error. At 466 taps the same residual is a full result page.
The gap between 0.60 and 10.0 is worth understanding before designing anything: the 77-tap
measurement used 50 natural-language queries averaged, and an average hides the shape here.
Duplicate pressure was already known to be a step function of which registries are tapped
rather than how many; near-identical pressure looks like a step function of which query —
harmless across a query set, total on any query that lands on a widely-mirrored skill. Re-measure
per-query maxima, not means.
The hard part is the safety proof, not the clustering. Content hashing was adoptable because
one count settled it: of 14,153 distinct bodies, clusters spanning more than one name numbered
zero, so collapsing could not merge two different skills. Near-identical clustering has no
such free proof — any similarity threshold loose enough to merge a Japanese translation with
its English original is loose enough to merge two genuinely different skills that share boilerplate.
Establish the equivalent bound first (over a real corpus, at the chosen threshold, count clusters
spanning more than one meaning) or the fix trades a visible problem for a silent one.
Three things to get right. Translations are the motivating case and the hardest: they
share almost no tokens with the original, so token-overlap similarity will not find them while an
embedding will — and the vectors are already on disk, which makes this cheaper here than it
would be anywhere else. Collapse before k, and at both the
retrieve and retrieve_any seams, for the reason the shipped dedup already
documents: fusion reintroduces copies either engine dropped, because the copies are distinct
(tap, skill_md) keys and RRF has no reason to treat them as one. The existing
quality prior carries over unchanged — rag.source_rank orders on the user's
curated flag first and shipped confidence second, and choosing among
near-identical copies is the same question as choosing among identical ones: where should the user
install from.
Not to be confused with #629, which deduplicated vector storage (one
row per distinct embedding, 39.7% repeats reclaimed). That is a disk-size fix beneath the index and
changes no ranking; this is about which rows reach the user's screen.
What shipped, and what did not. rag.collapse_near_duplicate_hits is the same
"keep the earliest rank slot, promote a better source" contract as dedupe_by_content,
run over cosine similarity of the entries' first-chunk embeddings
(dense.entry_vectors, an index probe through chunks_entry on a quantized
store) instead of a body hash, at the retrieve_any seam before k is
applied. It is covered by unit tests down to the arithmetic (_cosine's dimension-
mismatch and zero-vector guards), the clustering contract (rank order, quality-prior promotion,
limit-after-collapse), the dense.entry_vectors lookup against a real quantized
sqlite-vec store, and the retrieve_any/boost search
--collapse-near-duplicates wiring in both directions (on and off).
It ships opt-in and off by default — retrieve_any(..., collapse_near_duplicates=True)
or boost search --collapse-near-duplicates — rather than replacing
dedupe_by_content's output on the default path. Two things this card asks for are still
open, and both need a real embedding backend (a built dense index, over a real multi-tap corpus)
that the environment this was implemented in cannot reach — no network path to an embeddings
provider or to the local ONNX model download, confirmed rather than assumed: huggingface.co
and pypi.org both refuse at the network policy layer. First, the safety proof
this card itself demands before defaulting the mechanism on — “over a real corpus, at
the chosen threshold, count clusters spanning more than one meaning” — has not been run;
NEAR_DUPLICATE_THRESHOLD = 0.97 is a starting point, not a validated floor. Second,
re-measuring the exa search case (and per-query maxima generally) against the fix needs
that same corpus and index. Whoever runs that measurement should flip the CLI flag's default, fold
the corpus count into this card's evidence, and only then consider this shipped.
The bound has now been measured, and it says the acceptance test in this card is the wrong
one. scripts/measure_near_duplicate_bound.py runs the count this card asks for
against the pinned 20-repo eval corpus (10,152 entries, 104,271 chunks, BAAI/bge-small-en-v1.5
at 384-d). Those entries reduce to 5,714 distinct chunk-0 vectors — 44% of entries
already share a chunk-0 embedding byte for byte — and at
NEAR_DUPLICATE_THRESHOLD = 0.97, 162 pairs clear the threshold and 56 clusters span
more than one name. Sweeping the threshold moves that number but never to zero: 0.96 → 91,
0.97 → 56, 0.98 → 28, 0.99 → 13, 0.995 → 8, 0.999 → 4.
Four of those 56 are not the threshold's doing at all. They are clusters of a single vector
shared by several names, so they cluster at any threshold, which is why the sweep bottoms
out at 4 rather than 0. The largest is the same at every threshold and is worth naming: 28
differently-named agents from one tap (affaan-m/ECC — architect,
code-reviewer, chief-of-staff, database-reviewer,
e2e-runner, …) whose chunk 0 is the same Spanish preamble
(No cambiar rol, persona ni identidad…) in every file. Chunk 0 is
name + description + opening of body, and where a registry opens every file with
identical boilerplate, the name does not move the vector enough to separate them. A floor exists
that no threshold can reach under, so “count must be zero” was never achievable.
Worse for the test: most of the other 52 are the feature working. Hand-classifying all 56 at
0.97, roughly two-thirds are genuinely one skill under two names — twelve are pure
hyphen-versus-underscore renderings of one integration (zoho-mail /
zoho_mail, google_maps / google-maps,
anthropic_administrator / anthropic-administrator), and the rest are
suffix variants of one document (tdd / tdd-guide,
rust-review / rust-reviewer, testing-patterns /
code-showcase-testing-patterns). Collapsing those is precisely what this card exists to
do. A metric that counts them as violations would reject every threshold that works.
The dangerous merges have a shape, and this card already named it. The ~20 clusters that are
real false merges are dominated by near-miss brand names: coinmarketcal with
coinmarketcap, bugbug with bugsnag, parsehub
with parseur, linkhut with linkup,
mx-technologies with mx-toolbox,
salesforce-marketing-cloud with salesforce-service-cloud. These are
distinct products whose descriptions are boilerplate around a swapped word. That is the
core/typosquat.py confusion shape this card's opening paragraph pointed at, arrived at
independently from the other end: the guard this needs is not a tighter cosine floor but a
name-confusability veto — refuse to collapse two entries whose names are a confusable
edit apart, however close their vectors sit.
So the default stays off, for a better-supported reason than before. The measurement does not
say 0.97 is too loose; it says similarity alone cannot separate tdd/tdd-guide
(collapse) from coinmarketcal/coinmarketcap (never collapse), because both
pairs sit in the same cosine band. Flipping the default needs the confusability veto first, and a
re-count with it applied. And this bound is space-specific: it was measured in
bge-small 384-d, while a keyed production install is voyage-4 at 1024-d.
Cosine thresholds do not transfer between embedding spaces — rerun the script against each
space before trusting a number in it.
The monthly corpus refresh rewrote the pins and the baseline but left every documented number stale — taps.txt now contradicts its own header, and nothing checks it
With a corrupt config.json, doctor reports "no registries tapped" and verdicts "● ready to set up" exit 0 while search is dead; heal says "nothing to heal"
BOOST_NO_EMBED has no state in the reason ladder: doctor calls a deliberate kill switch a degraded fault (exit 1) and hands advice that is a measured no-op in both branches
status() has no state for "ready but the embedder does not work": doctor green-ticks a tier that never ran, the search hint is suppressed, and every search re-pays the failed model fetch
out.err's multi-line hint is coloured as one span, so line 1 ends with no RESET and lines 2+ carry no start code
The corpus's 65.6% duplication is 99.93% inside a single tap, so the required gate never once exercises the cross-tap trust ordering dedup exists for — 0 swaps in 264,735 comparisons
Measured. Over the 91 required golden queries on the current pins, 264,544 of 264,735 collapse comparisons inside rag.dedupe_by_content (99.93%) were between two copies in the SAME tap and the source-preference branch executed 0 times; and forcing it to fire — marking one tap curated produces 191 swaps — leaves recall@k / hit@1 / MRR / nDCG@k at 0.8407 / 0.4835 / 0.6065 / 0.6552 both before and after, with all 91 graded ranked-key lists identical.
Reproduce it.
cd <repo>
export HOME=$TMPDIR/audit-corpus-verify && export BOOST_HOME=$HOME/.boost # after eval_corpus.py --ensure (see other finding)
.venv/bin/python - <<'PY'
import sys, json
sys.path.insert(0,"."); sys.path.insert(0,"scripts")
from boost_cli.core import rag
from boost_cli.core.rag import source_rank
stats={"collapses":0,"same_tap":0,"diff_tap":0,"swaps":0,"ties":0}
def patched(hits, limit):
best={}; out=[]
for hit in hits:
d=hit.get("content")
if not d: out.append(hit); continue
s=best.get(d)
if s is None: best[d]=len(out); out.append(hit); continue
stats["collapses"]+=1; kept=out[s]
…
What adversarial verification corrected. This card's numbers are the re-measured ones, not the ones first reported.
1. THE REMEDY CLAIM IS WRONG, and this is the correction that matters. The finding's why_it_matters says "adding a single mirror registry to taps.txt would be enough to make it visible." It would not. eval_retrieval.grade_key (scripts/eval_retrieval.py:197-209) keys every ranked slot on the content digest (body:<digest>), the exemplar class (cls:...), or the entry name — and a swap only ever replaces hit["entry"] INSIDE a content cluster, where the digest is identical by construction and, because catalog._content_digest hashes name + description + body, so is the name. Every return branch of grade_key is therefore invariant under a swap; the nohash:tap::skill_md branch at :209 is unreachable for a swapped hit because dedupe never collapses a hit with no digest. Proven, not argued: marking composio-community/awesome-codex-skills curated on the same pinned corpus fires 191 swaps (every cross-tap collapse becomes a swap — the composio copy arrives second in all 191) and leaves recall@k / hit@1 / MRR / nDCG@k at 0.840659 / 0.483516 / 0.606517 / 0.655233 before AND after, with all 91 graded ranked-key lists byte-identical. So a trust-ordering regression is invisible to the required gate BY CONSTRUCTION, regardless of corpus shape — the gap is that the harness never grades on source, not that the corpus lacks mirrors.
What is NOT established. Recorded because a card that overstates its own evidence is worse than no card.
SCOPE OF MY REPRO — read this before re-verifying. The finder's numbers reproduce ONLY on the CURRENT pins (10,731 entries). The shared read-only corpus at $TMPDIR/eval-home is PRE-REFRESH: 10 of its 20 clones are not at the SHAs in tests/eval/taps.txt (anthropics/skills, NeoLabHQ, LessUp, affaan-m/ECC, first-fluke, langchain-ai, minio, OneWave-AI, sickn33, anthropics/claude-agent-sdk-python), and it holds 10,152 entries. There the same instrumentation gives 1,986 clusters / 6,348 entries (62.5%) / 240,646 collapses / 240,455 same_tap / 191 diff_tap / 0 swaps. I materialised the current pins into a private copy (eval_corpus.py --ensure, network fetch works despite a harmless failed to store: 100001 commit-graph warning) to get the finder's exact figures. Anyone re-checking against $TMPDIR/eval-home will get the smaller set and should not read that as refuting the finding — the qualitative result (5 cross-tap clusters, all high/high, 0 swaps, ~99.9% same-tap) is identical on both.
DOC TRAP, not this finding's fault: tests/eval/taps.txt's own header prose says "10,152 entries" and "sickn33 … is 6,309", and CLAUDE.md repeats 10,152 — but the file's own per-repo rows sum to 10,731 with sickn33 at 6,634. The pins were moved in commit cbc0a58b ("test(eval): refresh the pinned retrieval corpus") and the header prose was not updated. A card author quoting corpus size must take the row sum, not the header.
NOT ALREADY CARDED, but read the existing card first.
Why it is worth doing. The required corpus contains essentially none of the duplicate shape that dominates a real install. A user's duplicates arrive as mirror registries republishing each other's skills across taps, which is what source_rank decides between and what determines where a user is told to install from; the gate's duplicates are one publisher re-vendoring itself into 60 plugin bundles, where every candidate has the same tap and the tie-break is a no-op.
Found by an automated audit of retrieval/eval quality, the search & browse surfaces, and first-run onboarding; every finding was then re-measured from scratch by an independent adversarial verifier whose instruction was to refute it. Verdict: CORRECTED. No fix is prescribed here — the measurement is the contribution.
make eval scores a corpus with every SKILL.md body missing, reports all four floors PASS, and scores HIGHER than the real corpus
boost explain's heuristic fallback prints every heading in the file — 541 lines for one skill — while the sibling list in the same function caps at 12
The CI-vs-Makefile floor-parity test compares only the floor VALUES, so changing -k in ci.yml turns a PASS into a FAIL with the test still green
boost info/deps on a not-installed rule or workflow widens the tap's sparse cone for a directory source_dir_for immediately rejects
boost_install's description enumerates four of the five enabled agent targets and omits Antigravity CLI, the fourth linking agent
boost_search advertises "10-15 seconds" unconditionally; with no AI configured it is 0.013 s median and the rerank never runs
Write-up · mcp-search-cost-overstated-on-keyless-machines.md
The skip list — the one bound on boost's triggers — ships in INSTRUCTIONS only, in zero of the seven tool descriptions
out.panel still fits its content to term_width(), so boost count | … clips the line to an assumed 80 columns
quickstart silently discards the incompatible shard status, so a user with any API key is never told why zero vectors arrived
A local manifest read error is reported as "no published shards", and the BoostError's hint — the only actionable line — is discarded
Write-up · quickstart-manifest-error-drops-hint-and-misnames-cause.md
With every registry unreachable, quickstart prints "✓ indexed 0 items" and "✓ ready", exits 0 — and the command it recommends exits 1
Write-up · quickstart-says-ready-exit-0-after-every-tap-failed.md
The shard download is invisible both before and during: --catalog --dry-run never names the 1,604.8 MB, and the live fetch passes no progress callback and has no spinner
Write-up · quickstart-shard-download-invisible-in-preview-and-run.md
README's "81 commands" table enumerates only 80 — the missing one is quickstart, the README's own first command
Write-up · readme-81-command-table-lists-80-omits-quickstart.md
README's hero block ends in exit 1: boost install tdd-workflow names a skill no starter registry ships
Write-up · readme-hero-installs-a-name-no-starter-registry-ships.md
boost recommend sizes the description cell per row from that row's because: text, so neither column lines up
The corpus refresh re-baselines only golden.jsonl, so golden-natural.jsonl's baseline silently describes a corpus that no longer exists
The two roadmap boards' install footers still advertise Python 3.9+, four minor versions under the real floor
search caps the name column at 32 and the tap column at 20 at every terminal width, while the description column grows without limit
The search footer is the one unfitted line on a screen search_layout just fitted: 55 columns in a 40-column pane
Write-up · search-footer-not-fitted-to-the-pane-its-rows-were-fitted-to.md
boost search's footer overflows a 40-column pane by 2-15 cells, and the gate that swears it doesn't passes only because its fixture returns exactly one match
Write-up · search-footer-overflows-40-col-pane-gate-passes-by-luck.md
The relevance meter and its "one gradient moment" are constant on the default result page: 138/150 rows full bars, 150/150 the same colour
Write-up · search-relevance-meter-is-constant-on-default-page.md
boost search rows fit to term_width(), not pane_width(), so a piped search silently loses the TAP column entirely
Write-up · search-rows-use-term-width-so-piping-drops-the-tap-column.md
The contributor-onboarding gate tables state three wrong suite sizes, and README and CONTRIBUTING disagree with each other
rag.surface's de-hyphenated name copy is justified by two claims that are both false, and its real effect — an undocumented 3x name / 2x description field weight — is guarded by a test …
_fit_widths shrinks data columns to a bare "…" and still overflows: boost taps is 54 columns wide on every terminal narrower than 54
Write-up · table-fit-widths-floor-destroys-columns-and-still-overflows.md
Tier 3's false-call ceiling is unreachable at its own default N, and tolerates zero false calls at the N make eval-tools uses
An unreadable tap cache is invisible to doctor ("✓ 1 tap cloned & cached", exit 0) while search, browse and heal all exit 70 with a crash report
Write-up · unreadable-tap-cache-healthy-doctor-crashing-heal.md
"! agent dir ~/.cursor/skills is not writable" is the one doctor issue with no next action — heal has no path for it and the next install crashes at exit 70
Nothing boost prints ever names an entry point: bare ./boost is byte-identical on a virgin machine and a working one, and the one command its failure-hints route you to is the only setup …
golden-natural.jsonl — the only fully exemplar-graded query set — is invoked by no make target and no workflow, so its numbers can only be produced by a human typing the command
The "taps last refreshed N days ago" hint can never fire on a machine that tapped and never ran boost update — the only writer of the marker is update itself
CLAUDE.md instructs the wrong licence: every new source file is told to carry GPL-3.0-only in a repo whose 373 headers, LICENSE and pyproject all say Apache-2.0
At a narrow pane boost hooks list drops name, the argument hooks remove -n takes, while host and scope survive
An unwritable agent rules/ or commands/ dir still crashes a rule or workflow install at exit 70
Write-up · unwritable-rule-or-workflow-dir-crashes-install.md
Two config.json shapes still slip past the corrupt-config handling. Invalid UTF-8 crashes every command, doctor included. {"taps": "x"} is a fresh install to doctor and a configured machine to --help.
Two cache writers still crash on what one sudo boost leaves behind. A read-only _names.txt fails heal, update and untap at exit 70 while doctor says healthy.
Write-up · cache-writers-that-still-crash-on-a-read-only-cache.md
Just before out.table drops a column, it shrinks the widest one with an ellipsis. In hooks list that is often name, and bmad-r… is not a name hooks remove -n accepts.
Write-up · identifier-columns-shrink-to-a-name-no-command-accepts.md
Under a read-only ~/.boost with no cache dir, update, heal and doctor crash at exit 70, and heal --dry-run says 0 for a run that crashes
With an API key exported and no vector store yet, doctor and search send the user to a paid build while quickstart offers the free download
Write-up · keyed-no-store-hint-disagrees-with-shard-remedy.md
A rule or workflow row for a disabled agent makes boost sync claim the same repair on every run, and doctor never goes healthy
Agent-dir shapes boost still trips on: a parent with no search bit, a missing dir under a read-only parent, a dir at a rule's file path, a blocked agent dropped from the lock's scope
Found by the third review of read-only-boost-home-with-no-cache-dir. None is a regression.
Each was measured behaving the same on c7dca95c.
A parent with no search bit. With an agent dir's parent at mode 0o600,
store.install copies the skill into the store. It then crashes at exit 70 in
linked_agents, where (adir / name).is_symlink() raises
PermissionError outside link_agents' guard. The store dir is left
unrecorded. Treating an OSError there as "not linked" would close it.
A missing skills dir under a read-only parent. The install names a chmod
target that does not exist, and doctor and heal disagree about it.
link_agents' PermissionError branch should record the block from
paths.refuses_writes(adir) and word it through link_refusal. doctor's
agent-dir check should use the same function, so that it agrees with heal.
A blocked agent drops out of the scope. After an install skips an agent as blocked, the lock's
agents list leaves it out. preserved_agent_scope replays that list, so a
later install --force never retries the agent and never says so. Only
boost sync reports and repairs it.
Callers that drop the result. boost import (install_from_path),
quarantine --release and the unsideline paths (focus, profile, context, team) drop
res.blocked without a word, as they already drop res.unwritable.
pkg._warn_unwritable(res) is the existing reporter.
Also found by the reconcile review of #933 with #931, and also measured on both parents: With any agent dotdir at mode 600, on Python 3.12 and 3.13 (3.14's Path.exists answers False where they raise), doctor, heal and sync
exit 70 (PermissionError on ~/.cursor/skills). A rule or workflow install in that
shape is fixed on #933's branch: _refused_target names the dir through
paths.refuses_writes, which asks lexists. So is uninstall of a rule or workflow there, through store.refusing_dir, which now treats a path it may not look at as not there. The skill install and these three are not fixed. And the remedy is wrong for this shape: chmod u+w leaves a 600 dir at 600, because the missing bit is search, so the wording should say chmod u+wx when X_OK is what fails.
A directory at a rule's target file path (~/.cursor/rules/house.mdc/) raises
IsADirectoryError. That is not one of the refusal shapes, so install exits 70 after writing
the other agents' copies, with no lock entry.heal --dry-run previews "would re-materialize" a rule whose target dir is still locked or
blocked. The real run then re-materializes nothing, correctly. The exit codes agree (1 and 1), but the
wording does not.
With ~/.agents/skills read-only, every install exits 1, naming it, while
doctor says healthy. doctor could ask paths.refuses_writes(paths.store_dir()).
Six install paths still printed success over an agent dir that refused the link
A reinstall into a locked skills dir says "not linked" over a link that is already there and correct
Write-up · a-correct-link-in-a-locked-dir-reads-as-refused.md
A refusal and its remedy print past the pane, on the one line a user reads before acting
The import-budget gate reports OK when the command it measures never ran
Write-up · import-budget-gate-passes-when-boost-cannot-import.md
The release guard fails open when it cannot read the commit's tags
Write-up · release-guard-fails-open-when-it-cannot-read-tags.md
One failed shard-build job silently drops its whole chunk from the published manifest
Write-up · one-failed-shard-job-drops-registries-from-the-manifest.md
The CI job summary reports the coverage gate as 80% when it is 90%
One valid-JSON message that is not an object kills the whole MCP session
Write-up · mcp-server-dies-on-a-valid-json-message-that-is-not-an-object.md
boost_doctor certifies a machine healthy that boost_search calls untapped
Write-up · mcp-doctor-says-healthy-where-mcp-search-says-nothing-is-tapped.md
MANIFEST.in prunes only the root copy, so dev-local files ship in the sdist
A search that finds nothing echoes the whole query back into the agent's context
A deeply nested JSON line kills the MCP session, because the parse guard names one exception
The model back-off dated its record by the filesystem clock and read it against the process clock
boost mcp registration writes outside a sandboxed HOME
Write-up · mcp-registration-writes-outside-a-sandboxed-home.md
Codex hooks are shaped like Claude's, and the two fields that differ are unverified
Registering a skill's MCP server handed Antigravity the one argv its CLI rejects
Found by a reviewing subagent on #977,
which was about a different file. boost install of a skill that declares an MCP
server offers to register it with every agent CLI on PATH, and built that command line in
core/mcpdecl.py — a second copy of the per-host grammar whose docstring said it
"mirrors core/mcphost.py exactly". It mirrored two hosts of three. It had a Gemini
branch and a Claude fallthrough, and agy fell through to Claude's — the single host whose CLI
rejects that shape.
Measured against the old code: agy mcp add gh --scope user -e K=v -- npx -y gh-mcp.
Antigravity CLI has no scopes (one global file at
~/.gemini/config/mcp_config.json, inherited from Gemini CLI), so --scope is
an error rather than a no-op, and it requires every flag before the name, so the
-e is rejected too. Every declared server, every install, on any machine with
Antigravity installed — reported as "could not register gh" and never once as a boost bug, because
the line the user sees is the CLI's.
The fix is not the missing branch. Adding one would have left the same two copies, with the
fourth host (Codex, next) to remember in both. The grammar moved into one function,
mcphost.add_argv(host, name, command, tail, …), and both callers became thin:
register_argv registers boost itself (launcher mcp --stdio) and
mcpdecl.register_argv turns a declared spec into the same call. Output is byte-identical
for Claude and Gemini — all 134 existing argv assertions pass untouched — and correct for agy for the
first time.
And the fallthrough is now loud. Claude's shape was the return at the bottom of
the function, which is why a host with no branch inherited it silently. It is an explicit
if host == CLAUDE, and the bottom of the function raises. A host can be in
HOSTS and have no grammar for exactly as long as it takes a test to run.
Then a verifying subagent found a worse bug in the code that was now shared — and it was
Gemini's, the host the old copy had "mirrored" correctly. boost emitted no -- for
gemini at all, on the belief that its trailing variadic would pass the separator through as a
literal argument. Measured against the real Gemini CLI 0.61.0: it does not.
unknown-options-as-args rescues only options gemini does not know, and
gemini mcp add knows ten of them, so the canonical GitHub server spec — which ships
the bare docker run -i --rm -e GITHUB_PERSONAL_ACCESS_TOKEN ghcr.io/… — had its
-e claimed by gemini's own --env (nargs: 1, taking the
variable's name as its one value) and dropped from the stored entry — exit 0, "server
added", and a container launched with no token. The -e KEY=value spelling is not
dropped but relocated, into gemini's own env map and out of the args docker
reads, which is the same outcome by a different route and the one to expect when reading a real
entry. A -t http in a spec's args is worse: the entry is rewritten as
{"url": "npx", "type": "http"}. The separator goes after the command for
gemini — the one host where it does — and it is emitted unconditionally, because
add x npx -- is accepted and stores args: [].
The same pass killed the comment justifying agy's --: measured on 1.1.22, agy never
consumes an argument that follows the command, and rejects a dash-leading command with or without
it. The -- stays (it is agy's documented separator) but is now described as inert
rather than load-bearing. And boost install's decline path printed one command
line for targets[0] after a prompt that named every host — so a user who said no got
Claude's line and nothing for the two hosts a yes would also have registered with, agy among them.
Fifteen new tests and two rewritten ones that had pinned the old, wrong shape. The three literal
per-host argvs on the install path (none existed: nothing tested mcpdecl.register_argv
with a host= at all); agy's two rules each pinned separately, so a regression names
which one broke; gemini's separator pinned three ways, including a spec whose args start with a dash
— the case that was silently losing data; a cross-module parity test parametrised over
mcphost.hosts(), which is the docstring's promise turned into an assertion and which
covers a fourth host the moment the table gains one; and a monkeypatched codex row
asserting ValueError — the regression this shape exists to prevent, failing in CI
instead of on a user's machine.
wrap() sizes a stderr line by stdout's pane
A line now gets its colour from the stream it lands on and its
width from somewhere else. _wrap_lines
(boost_cli/core/output.py:235) folds to
term_width() - lead, and term_width
(:402) takes no stream: it calls
shutil.get_terminal_size, which consults
sys.__stdout__. Right beside it, pane_width(stream)
(:407) does take one, and exists because the same question has a
different answer per stream.
So the emitters that write off stdout and wrap — err(wrap=True)
(:318, :327) and warn(stream=sys.stderr,
wrap=True) (:284, reached from catalog.py:346,
journal.py:65, lockfile.py:185,
registry.py:594 and complete.py:105) — size a stderr
line by stdout's pane. In a 200-column terminal:
boost install x >out folds the hint on the terminal at 72
columns, because stdout is a file, get_terminal_size raises and
the 80-column fallback applies; boost install x 2>log writes
192-column lines into the log.
This is the twin of err-colour-judged-by-stdout, found while
verifying it, and deliberately left out of that PR: it is a width bug, not a
colour one, and it predates the change. Fix sketch: give
term_width (or _wrap_lines) the stream the emitter
is about to write to, the way c(), role() and
aurora() now take one, and have each wrapping emitter pass its
own. Watch the COLUMNS precedence: pane_width
honours an explicit COLUMNS either way, and the wrap path must
keep doing the same or a scripted COLUMNS=100 stops applying to
hints. A test wraps one long hint with stdout a TTY and stderr a plain
buffer, and asserts the fold point comes from stderr.
Measured before claiming, and the fix sketch above is wrong.
pane_width(stream) is not the per-stream model to copy: it uses
its stream only for the isatty() predicate and then takes the
number from term_width() (output.py:422),
so it measured 80 for a 200-column stderr terminal while stdout was a file.
Routing _wrap_lines through it would have left that intact. A
third instance the card missed: spin.progress_clear
(spin.py:98) erases term_width() columns of a line
it just proved is on s · 80 of a 200-column bar.
Two numbers here were budgets rather than widths. 192 is
term_width() - len(" hint: "); the lead is printed too, so the
lines run to the full pane — measured 196 for a hint and 199 for a warn
on a 200-column pty. And the proposed test cannot fail: a fake tty (a
StringIO whose isatty() returns True) has no file
descriptor, so broken and fixed both answer 80 — measured identical fold
points either way. It needs a real os.openpty() with
TIOCSWINSZ.
Engine & command internals
// concrete file:line findings from the code scanSemantic search is gated behind an API key it does not need
Every dense search re-scanned all 3.08 GB of vectors — vec0 has no ANN index
Memoize config.load() in-process
Atomic skill install (temp-dir swap)
One shared atomic-write helper
Unify _tilde() — two copies have a boundary bug
Cache the catalog entry-set across RAG queries
Write-up · cache-the-catalog-entry-set-across-rag-queries.md
Stop re-serializing entry meta on every search
Write-up · stop-re-serializing-entry-meta-on-every-search.md
Prune ignored dirs during scan_dir walk
Single tech-stack prober
Single imperative-rule extractor
Split oversized command modules
Robust tag argument parsing
Localize the stored BM25 snippet
Frontmatter scalar over-coercion
Rule install — materialize rules into each agent's native format
Workflow install — drop commands/subagents into each agent's native dir
Workspace scope — boost install --local into the project
boost list shows installed rules and workflows
Teach the rest of the CLI about project scope
boost update refreshes installed rules and workflows
Scan and sync rules/workflows like skills
boost install --scope user|project for rules/workflows
Ambiguous tap short-name resolution silently picks the wrong tap
sync --apply deletes any broken symlink, not just boost's own
Dense search's empty result skips the BM25 fallback
Journal rotation has a lost-update race between concurrent processes
boost uninstall has no confirmation prompt
update/reinstall silently widen a skill's agent scope
lint --tap mis-scores rule/workflow entries as broken
sync relinks a narrowed skill into every agent
Dependabot raises every toolchain bump twice
Write-up · dependabot-root-pip-entry-duplicates-requirements.md
boost onboard silently overwrites existing generated files
Write-up · onboard-overwrites-generated-files-without-confirm.md
Dependabot cannot regenerate the hash-pinned locks
Negative -n silently inverts log/pulse output
The toolchain lock has no proactive update path any more
boost log --crashes listing branch has no non-empty test
boost serve's own path-traversal guards are untested
Diagnostic log has no structured/JSON output mode
AI bridge swallows failures with zero diagnostic trail
The eval gate reports a perfect score on a corpus 4× smaller than a real user's
boost info rejects the tap-qualified name its own error tells you to type
The axe-core sweep intermittently fails color-contrast on a page the PR never touched
The eval gate would not pass on the catalogue its own users have
Write-up · eval-corpus-is-96x-smaller-than-a-real-install.md
The golden set grades by name, and 35 of 53 names are ambiguous
The “pinned” eval corpus pinned names, not commits
The eval de-duplicated its ranked list by name, so homonyms shared a rank
62% of the required gate's corpus is a single third-party repository
Nothing refreshes the eval corpus pins, so the gate measures one frozen day
agents recorded the request, not what was linked
Write-up · agents-field-records-the-request-not-the-links.md
the Python floor moves from 3.9 to 3.12
the typing.List → list sweep the floor now allows
audit the 16 zip() calls the 3.12 floor made checkable
boost search never noticed a tap added after the first search
boost update told you to run a flag it then rejected
one deleted upstream stopped boost update for every other tap
The update path skipped the scan the install path runs
heal removed symlinks boost never created
bring scripts/ under the ruff gate
install --dry-run predicted an install that never happens
boost_search never said which ranking produced its answer
est_items counted one skill fourteen times once registries went multi-agent
The catalog was missing the two most-starred token-efficiency registries
Dependabot splits one action repo into three unmergeable PRs
The shards workflow has never once produced a shard
The repair command could not repair the thing two commands sent you to it for
Taps download and check out the 84% of a repo that boost never opens
A catalog entry knew where it came from, never what it was
a convention that said "verify the repo is real" verified nothing
Write-up · the-catalogue-advertised-repos-that-no-longer-exist.md
Path.exists() looks total, and is not
a publisher that could not publish, and an alert that could not stand down
the shard job that had never once finished, and the timeout that could not be raised
the scheduled re-pin that refreshed the twenty rows it must not touch, and none of the hundred and sixty-five it existed to pin
--category marketing matched nothing, while four marketing registries sat in the catalog under other names
reindex --dense runs for hours behind a bare spinner, and a cancel discards all of it
Dense reuse is per tap, so one changed file re-embeds the whole registry
The dense store keeps one vector per copy, not one per distinct text
adapt renders every sibling of a flat agents/-dir workflow as a 138-agent crew
Write-up · audit-adapt-renders-every-sibling-of-a-flat-agents-dir-workflow-as.md
reinstall, sync repair and audit resolve by name+tap, ignoring the lock's source path
Write-up · audit-reinstall-sync-repair-and-audit-resolve-by-name-tap-ignoring.md
Pinned taps silently moved to HEAD by update re-clone and compact --reclone, pin left stale
Write-up · audit-pinned-taps-silently-moved-to-head-by-update-re-clone-and-co.md
focus/profile sideline by unlinking without recording it; list lies, doctor exits 1, and its own remedy (sync) undoes the switch
Write-up · audit-focus-profile-sideline-by-unlinking-without-recording-it-lis.md
snapshot restore replaces the lock wholesale, orphaning newer rules' CLAUDE.md blocks
Write-up · audit-snapshot-restore-replaces-the-lock-wholesale-orphaning-newer.md
boost create: CLI audit findings (2026-08)
boost reindex: CLI audit findings (2026-08)
quickstart reruns re-download every shard; shards.sync never asks what is built
Two concurrent rag.build() runs delete each other's temp index
Write-up · concurrent-rag-builds-delete-each-others-temp-index.md
compact --reclone left a clone on HEAD when it could not reach the pin
Write-up · reclone-leaves-a-clone-on-head-when-it-cannot-reach-the-pin.md
compact <tap> answers a question about one tap with a global all-clear
Write-up · compact-answers-a-named-tap-with-a-global-all-clear.md
boost taps vouches for a clone that is not there
reindex reported a completeness share it could not compute
Write-up · reindex-reported-a-completeness-share-it-could-not-compute.md
A scan that stops the bare-int count flag coming back
The sandbox fixture sandboxes every env var except the one project scope reads
`tests/conftest.py`'s `sandbox` fixture is thorough about the environment — it redirects `HOME`, clears `BOOST_HOME`, `BOOST_AGENTS_STORE` and `CODEX_HOME`, and sets five `BOOST_NO_*` guards so no test reaches AI, the network, the seed or a confirm prompt by accident. It never touches the **working directory**, and that is the one input project scope resolves against: `scopes.resolve_base` walks up from `os.getcwd()`, so a test that installs with `scope="project"` and no explicit `base`, or drives `boost install --local` without `monkeypatch.chdir`, writes into the developer's own checkout. What lands there is not a stray temp file. A project install materializes the full agent fan-out at the repo root — `.boost/skill-lock.json`, `.claude/skills/`, `.cursor/`, `.windsurf/`, `.gemini/`, `.codex/` — plus, for a rule, the context files themselves: `AGENTS.md` and `CLAUDE.local.md`, both of which every agent working the repo then loads as instructions. It is also self-propagating, because the state outlives the run that wrote it. A killed run left a project-scope `brainstorming` at the root; the next full suite read it back and failed twelve tests across four files that assert an empty install state — the `test_cli_quality.py` case for a machine that never installed anything saw `brainstorming ok project`, and `boost list`, `boost info` and the pane-width tests all inherited a row nothing in them had installed. The failures name none of the tests responsible, and they clear only when someone thinks to look at `git status` — where `.boost/` is gitignored and does not appear. No committed test writes there in its shipped form: `tests/unit`, `tests/functional` and `tests/smoke.sh` each leave the checkout clean when run directly, watched by a teardown hook that names the offending nodeid. A full `make check` does not — it ends with `.boost/skill-lock.json` and five agent dotdirs at the repo root, all stamped inside the mutation stage. That stage runs those same tests against mutated `boost_cli/core`, so the blast radius of a mutant that weakens a scope guard is the developer's own working tree, and 4,343 of them survive each run. Which is the whole argument for the guard: the suite is one edit — or one mutant — away from writing here, and when it does, the failures it causes name everything except their cause. Two halves, both cheap: - **Chdir the sandbox.** Point `cwd` at a directory under `tmp_path` in the fixture, so a forgotten `chdir` resolves to the sandbox instead of the repo. Needs an audit first — any test that reads a repo-relative path (the generated-file `--check` gates, `test_spdx_headers.py`) must take the repo root explicitly rather than inheriting it. - **Fail loudly if it happens anyway.** A session-scoped autouse check that snapshots the known project-scope artifact names at the repo root and fails the run — naming the test — when one appears. That is the half that would have turned twelve misleading failures into one accurate one.
The links job fetches 467 URLs from one host at once and GitHub 503s a random handful
Capping the burst did not stop the 503s — the whole fetch was signal-free
One 30-minute cap for nine matrix cells, and only Windows is anywhere near it
Write-up · the-windows-test-cap-is-five-minutes-from-the-measurement.md
Code health & security
// planned · free tooling to catch vulns, smells & bugsSecurity linting — bandit via ruff S rules
Dependency CVE gate — pip-audit
Retrieval-quality eval harness (Tier 1 + rerank lift)
Supply-chain posture — OpenSSF Scorecard
Grow & diversify the golden set for statistical power
OpenSSF Best Practices — all 67 passing criteria answered
Secret scanning — gitleaks + push protection
Complexity & dead-code radar — xenon · vulture
OSPS Baseline — three levels, audited rather than assumed
Significance-tracked engine comparison (ranx monitor)
Quality dashboard — SonarCloud (free for OSS)
OpenSSF silver — reachable solo, and mostly already true
Property-based tests — hypothesis on the parsers
Write-up · property-based-tests-hypothesis-on-the-parsers.md
Coverage-guided fuzzing — atheris / OSS-Fuzz
OpenSSF gold — how far it goes without a second human
Fork-safe network layer — explicit ProxyHandler
The README read like a machine wrote it — measurably
main went red because "require branches up to date" was never actually on
Log timestamps are local time mislabeled Z
mypy's default mode skips untyped command bodies
ruff 0.16 widens the default rule set — 83 new errors on a version bump
The lock invariant can't parse name[extra]==version, so a valid pin fails the gate
boost doctor checks installed rules and workflows
rel_time tests race the wall clock and flake on a loaded runner
Gemini logged a skill conflict every session and boost doctor called the machine healthy
A machine-readable VEX feed, sourced from findings that already existed
A missing or corrupt lock file makes doctor prescribe boost sync/boost heal, and both delete every live agent symlink of an intact install
Write-up · doctor-prescribes-sync-that-deletes-live-links.md
import --all skips the injection/secret scans and the per-skill report single import runs
Write-up · audit-import-all-skips-the-injection-secret-scans-and-the-per-skil.md
verify/drift say 'nothing installed' (exit 0) when the lock file is missing or corrupt
Write-up · audit-verify-drift-say-nothing-installed-exit-0-and-doctor-says-lo.md
Quarantine/pin state invisible: list/info/doctor report links and materializations that were removed
Write-up · audit-quarantine-pin-state-invisible-list-info-doctor-report-links.md
boost audit: CLI audit findings (2026-08)
boost doctor: CLI audit findings (2026-08)
boost drift: CLI audit findings (2026-08)
boost fingerprint: CLI audit findings (2026-08)
boost health: CLI audit findings (2026-08)
boost lint: CLI audit findings (2026-08)
boost policy: CLI audit findings (2026-08)
boost quarantine --release: CLI audit findings (2026-08)
boost test: CLI audit findings (2026-08)
boost verify: CLI audit findings (2026-08)
Three commands that denied what was on the disk in front of them
The lock kept advertising links it had just removed
One unresolvable agent dir silences the whole `boost doctor` report
Found by the review of loop/codex-agent-target.
paths.expand now raises a BoostError for a ${VAR} with no fallback
and nothing in the environment, rather than resolving it to an empty or literal path — the change that
stopped a mistyped agents.<name>.dir silently installing into
./skills. agents.known_agents calls it for every configured agent, so one bad
row now aborts every command that asks who the agents are: doctor,
heal, clean, install, sync.
For most of those, refusing is the right answer. For doctor it is not: its whole job is to
surface a misconfiguration, and it exits 1 having reported nothing else — no lock check, no store check,
no duplicate-discovery check. The remedy is reachable (boost config set never calls
known_agents, verified in a sandbox), so this is a bad report rather than a lockout.
Fix direction. Catch BoostError around the agent lookups in cmd_doctor
and render it as an issue row, then continue with an empty agent set. The catch has to cover the
indirect callers too — store.duplicate_discovery, store.sync_plan and the
per-agent coverage counts all reach known_agents — so the natural shape is one resolution
step at the top of the command whose failure degrades the agent-dependent sections rather than the run.
An item materialized nowhere reads ok in verify, health and drift
Found by the second review of sync-repairs-a-disabled-agents-row-every-run.
That item fixed a row for an agent boost no longer writes being read as a missing artifact:
nothing wrote the file and nothing ever will, so boost verify failed forever and
boost drift and boost health sent the user to a boost sync
that skips the row by design. The fix is right, and it opens a gap at the other end.
Skipping a row is correct per row and wrong per item. Skip every row an item has
and the item is materialized nowhere — the rule reaches no agent, the workflow is in no command
palette — and the surfaces say it is fine. Measured on a real install (fixture tap, one rule,
then every agent disabled): boost verify →
team-conventions ok rule, rc 0 · boost drift
→ in-sync · boost health →
rules 1 installed, drift 1 in-sync,
● healthy · boost attest --verify → sha_ok.
doctor both names it and denies it. It prints the note —
5 recorded materializations name an agent boost no longer writes… — and two lines later
✓ 1 rule and 0 workflows fully materialized for every agent boost writes, then
● healthy. The count is true as worded and false as read: the write set is empty, so
every item is trivially materialized for all of it. Only the note carries the news, and only
doctor has a note channel — a third state that is neither an issue nor silence. The
other commands have two, so the per-row skip had nowhere to put this and it fell into the healthy
one.
Why it is not cosmetic. verify is what scripts gate on and
attest --verify is the provenance check; an install that reaches nothing and passes
both is wrong in the direction that costs the user something — they installed a rule believing
some agent would read it.
Shape of a fix. Keep the per-row skip; add a per-item question after the loop. An
item whose materialization rows are all unwritten is reported distinctly — unreachable,
neither ok nor missing — and the count line says how many agents it
actually reaches rather than how many of the empty set it satisfies. The work is the state the
three commands do not have: verify's exit code, what health's counts
mean, and whether attest --verify should fail it. An item with no rows at all
is a different case (a skill, or a rule installed before rows existed) and must keep reading
ok.
Pipeline & supply-chain integrity
// planned · free tooling to secure the CI/CD path itselfPrebuilt vectors are published where no new user can reach them
Workflow SAST — zizmor
Every weekly shard run re-embedded the whole catalogue
Workflow linting — actionlint
Build provenance — SLSA attestations
SBOM on every release — CycloneDX / Syft
SBOM-aware scanning — osv-scanner
Second type checker — pyright
Widen the ruff rule surface — B·SIM·C4·PERF·RUF
Patch-coverage gate — diff-cover
publish.yml ignores the pip-audit / metadata gates
No timeout-minutes on any CI job
Dependabot's pip entry misses pyproject.toml's extras
License-compliance scanning of the dependency closure
The pytest tmpdir CVE — unfixable, then closed by the floor move
OpenSSF Scorecard's findings, triaged into three piles
main has no branch protection, so the release rules are honour-system
adapter-conformance's LangGraph leg never passed — a quoted matrix value
The required-check gate could not see paths:, so it green-lit a list that deadlocks every PR
Write-up · required-checks-can-declare-a-check-that-deadlocks-prs.md
ci.yml's job summary could exit 1 on its own, under the always() it was given
post-deploy.yml's always() destroyed the second signal it existed to preserve
markdownlint linted the fuzzer's corpus, so shipping a crash reproducer would redden a prose gate
Enabling a merge queue would deadlock every required check except ci.yml's
Write-up · merge-queue-would-deadlock-most-required-checks.md
The mutation gate was CI — 26 minutes, three times the next-longest job
The mutation gate's floor is a single file — shard 0 is store.py
A third of CI job time is spent waiting for a runner, not running
The required lint job pins zizmor==1.27.0 — a yanked release
The release verifies one commit and ships another
demo.yml still fails on every push — and the fix is not in the workflow
68 of the repo's 77 branches are merged loop/* branches nobody deletes
One commit can cut two releases, and the naive guard against it breaks retries
Near-duplicate items consume the top-10, and it gets worse with every tap
The BM25 index is one JSON blob, and it stops working between 10k and 50k items
Semantic search for users who will never set an API key
Dense retrieval today needs the [rag] extra and a
VOYAGE_API_KEY/OPENAI_API_KEY and a built
store. Most users will do none of that, so the default experience is BM25
forever. The keyless path is a local static embedding model —
potion-retrieval-32M class, MIT, model2vec family — which is not a
transformer: the entire weight file is one lookup table, so inference is
tokenize → gather rows → mean-pool → L2-normalize.
Measured locally, pure stdlib: ~1 ms to embed a query (mmap + bisect over
sorted keys), 12.8 MB of int8 vectors for 50k items at 256-d, and ~20 ms to
rerank BM25's top-200. No numpy, no sqlite-vec, no ANN index,
no new runtime dependency. import numpy alone costs 180–390 ms
cold, which disqualifies it from a one-shot CLI query path; the BM25 prefilter is
what makes the stdlib version viable, since a full 50k brute-force scan is 1.5 s
in pure Python. Pool depth is justified by measurement: BM25 recall saturates at
0.890 by depth 200 and gains nothing at 400, so reranking the top-200 gives up
essentially nothing versus scanning everything.
Why this and not a shipped Voyage index. A precomputed Voyage index is
inert without a Voyage query vector, and the only keyless way to get one
is a maintainer-run anonymous embedding endpoint — an unauthenticated free
embeddings API backed by the maintainer's card, which also ships every user query
off-machine and breaks offline. A local model is deterministic, so the artifact
becomes a cache rather than a correctness dependency: a newly tapped
repo can be embedded on the user's own machine. Doc-side is the asymmetry worth
shipping for — 29 ms/doc in pure Python is ~24 min for 50k single-core, versus
seconds in CI with numpy.
Do not ship this before the eval and dedup items. The headline claim
(+11.0 recall / +15.9 hit@1) did not survive verification: its baseline
used the kind oracle the real search path lacks, and both the blend weight
(w_dense=0.7) and the pool depth were argmax'd on the same 82
queries they were reported on, by 2-query margins. On a binary metric at n=82
the smallest net win reaching p<0.05 is 6 queries; hit@1 (+13 net
queries) holds up, recall (+9) sits at the resolution floor. And the structural
risk is real: the name is only ~10.5% of a mean-pooled surface vector while 106
description clusters are shared across 270 distinct names, so the lift may
shrink toward 50k rather than hold. Sequence: fix the gate, dedup, fix
the index format, then re-measure with McNemar and a held-out blend weight,
leading with hit@1. Test entry-level dense alone before the
blend, ship 512-d not 256-d (the whole case for 256-d was one query), keep
Voyage/OpenAI as the opt-in ceiling, and keep BM25 as the floor. The model table
cannot go in the default wheel — the shipped runtime is 0.79 MB and every merge
to main cuts a release, so +17.4 MB × ~24 releases/day exhausts PyPI's 10 GB
project quota in under a month; it needs a separate, rarely-released data
package behind an extra.
Related, and partly overtaken: [[keyless-semantic-search-for-everyone]] shipped a keyless
path using a transformer (BGE via ONNX Runtime, in the [rag] extra) while this
item was open. That does not settle the question this card asks — a static lookup table is
still far cheaper, and this card's discipline about not shipping a retrieval claim before the eval
still stands. What it does change is the baseline: "keyless" is no longer the differentiator, so the
case for a static model now rests on cost (~1 ms and no runtime dependency, against a
measured 233 ms cold and 34 MB of wheels) rather than on availability. Status left alone
deliberately — this is another loop's item to own.
Unblocked, and the case for it got stronger. This card says “do not ship this before the
eval and dedup items”. Both have now landed: the eval gate floors four metrics over a
realistic-sized corpus with baselines keyed to their query set, and content-hash dedup has merged.
The new evidence is a timing measurement. Building the shipped ONNX keyless store over 743
entries (3,740 chunks, bge-small-en-v1.5 on CPU) took 4,431 s — 74 minutes,
about 1.2 s per chunk. This card's static-embedding proposal claims ~29 ms/doc in pure Python. If
that holds it is a difference of more than an order of magnitude on the doc side, which is exactly
the cost that makes prebuilt shards mandatory today. Worth measuring the model2vec path directly
before committing — but the gap it claims to close is now a measured number rather than an
estimate.
Spike done — the prerequisites this card set have all shipped, so the measurement it asked
for was finally runnable. It says “do not ship before the eval and dedup items”;
dedup landed in #370, the index format in #367/#371, the
published eval in #373. What follows is potion-retrieval-32M (MIT, 63,091
× 512 F32 lookup table — confirmed a single tensor, no transformer) driven by a
hand-written pure-stdlib loader: WordPiece → gather rows → mean-pool → L2, mmap'd,
no numpy.
The two unverified performance claims were not just right, they were conservative. Query
embedding measured 0.16 ms against the card's ~1 ms. Document
embedding measured 1.34 ms on a synthetic 105-token doc and 3.27 ms on 300
real catalogue entries (median 61 tokens), against the card's ~29 ms.
That reverses one of this card's design arguments. The doc-side cost was the reason prebuilt
artifacts looked mandatory: “29 ms/doc in pure Python is ~24 min for 50k
single-core”. At the measured 3.27 ms it is 2.7 min — roughly the time
a first boost tap --defaults already takes. Local embedding is therefore viable on the
user's own machine, and shipped shards become a genuine optimisation rather than a requirement. (Not
to be confused with the ONNX bge-small path measured at ~1.2 s/chunk in
keyless-semantic-search-for-everyone; that number stands, and the gap between them
is the case for the static model.)
But reranking bought nothing at real scale, which is the result that matters. Over the 50
natural-language golden queries against a real 71,655-entry catalogue, reranking BM25's
top-200 by cosine scored hit@1 2/50 — identical to BM25's own
2/50, a net change of +0 queries where this card's own statistics note says 6 net
queries is the smallest win reaching p<0.05. On two hand-checked pairs the ordering was right but
the margin was thin (related 0.154 vs unrelated 0.097).
Stated limits, because this does not settle the question. The document vector was built from
name + description truncated to 1,500 characters, not the full body the real dense path
indexes, so this measures a weaker representation than the one being proposed. No blend was tried
— pure rerank, no w_dense — and this card explicitly asks for a held-out
blend weight and McNemar. What it does establish is that the cheap version of the idea does
not pay for itself, so the remaining work is representation and blending, not inference speed.
An unrelated finding fell out of it, and it is the more important one. BM25 scored
hit@1 0.040 here against the 0.340 published in #373. Both
are correct: the published figure is measured over the pinned 6-tap eval corpus of 743
entries, and this run used a real 77-tap install — 96× larger. Golden targets are
all present and rank 7th, 8th, 38th, 163rd rather than 1st. The eval corpus is not a scale model of
a real install, and the gate's floors describe a catalogue two orders of magnitude smaller than the
one users have. Tracked separately in [[eval-corpus-is-96x-smaller-than-a-real-install]].
Declined on measurement, after a second model was tried specifically to avoid declining on one
data point. The card's premise is that a local static model buys keyless semantic
search. Tested against the 50 natural-language golden queries over the pinned 20-tap corpus (3,843
entries as those registries stand today), with BM25 at hit@1 0.260 as the
baseline in every run:
potion-retrieval-32M (retrieval-tuned, 63,091×512 F32) scored dense 0.220,
hybrid RRF 0.260. potion-code-16M-v2 (code-domain, 63,457×256 F16 —
chosen because this catalogue is coding-agent skills, which is the strongest hypothesis for
why a general retrieval model would underperform here) scored dense 0.240, hybrid
0.260. A third representation — name+description only, over the full 71,655-entry
catalogue — reranked BM25's top-200 to +0 net queries.
So: two models, three representations, no measurable gain, and fusion never beats BM25 alone.
The code model is one query better than the retrieval model, which at n=50 is inside the noise
(±0.02 per query) and should not be read as a trend.
The contrast is what makes this a decline rather than a shrug. The published eval measures a
real embedding model at hybrid 0.440 against BM25 0.340 — a genuine
+0.100. Static embeddings reproduce the cost profile that made the keyless tier attractive
(0.16 ms/query, 3.3 ms/doc, no dependency) but not the quality that made it worth
having. Cheap and no better than what ships today is not a tier; it is a second code path to
maintain for nothing.
What survives. The performance findings stand on their own and are already recorded above:
doc-side embedding is ~24× faster than the card assumed, which is why prebuilt shards are an
optimisation rather than a requirement for the real models in
[[keyless-semantic-search-for-everyone]]. The pure-stdlib loader (WordPiece → gather →
mean-pool → L2, mmap'd, F32 and F16) is proven workable if a future model justifies it.
What would reopen this. A static model that actually separates on this task — the bar
is beating 0.260 as a reranker, not merely producing plausible cosines. The two hand
pairs looked fine for both models (0.224 related vs -0.001 unrelated for the code model), which is
exactly why plausible similarity was not accepted as evidence.
boost chat — ask about skills in plain language, grounded in retrieval
One global concurrency group let any PR cancel any other PR's check
The fuzzer found a real crash and nobody was listening
The published metrics could never be published
A line-anchored suppression drifted, and make lint never ran the tool
Write-up · zizmor-ignores-drift-and-make-lint-never-ran-it.md
The sweep gate died at browser launch, not at a page
A cold search materialises 71,600 entries to print five
boost search brainstorm finds nothing, and brainstorming finds it
Reproducible release builds — the sdist half nobody's setuptools does for you
The weekly republish reached the machines that had never been set up
A query made only of characters tokenize drops returns zero results, and the documented catalog.search fallback is unreachable — boost search "C++" finds nothing on a machine holding …
Write-up · bm25-empty-tokenization-kills-catalog-fallback.md
boost search 'C++' returns zero and blames the catalogue: tokenize drops every 1-char token, and nothing ever says a term was discarded
Mutation shards re-run every mutant on every push of the same PR
The sweep gate died at a page, not at launch — and took six unrun checks with it
A new card filed as shipped always 404s its own write-up link
Write-up · a-new-shipped-card-always-404s-its-own-write-up-link.md
Codex is a skills and rules target but not yet an MCP host
The free-threaded canary is red for a missing setuptools, so it warns about nothing
One 30-line class makes registry.py unsplittable, and floors the mutation gate at 95% of its timeout no matter how many shards you add
Found by measuring the first eight-shard run on main against its own prediction.
The planner said 28.4 min a shard; run 36959485751 measured 31.2, 39.6, 41.0, 41.0, 42.3,
43.1, 47.5 and 51.5 — a median of 41.0, which is 1.44x the prediction. Job
overhead is not the gap: the mutate shard step alone was 30.6–42.5 min.
The committed weights simply understate the work, and #1022 is the bot's re-fit that says
so (57,680,336 ms against 42,506,911, a ratio of 1.36).
The part that is a bug, not a drift. #1022 cannot merge: at eight shards its pack
scores a 73.5-minute tail against a 75-minute cap, which plan --timeout-minutes
correctly calls TOO TIGHT. The documented remedy is to re-pack — and re-packing does
nothing. At 10, 12 and 14 shards the planner returns the identical verdict, because
largest unit: 7309374 is registry.py whole, and one indivisible
unit is a floor on the slowest shard that no shard count divides. The only lever the repo
documents is connected to nothing.
Why it is indivisible. top_level_symbols returns [] for any
module holding a class with methods, because mutmut mangles a method's name differently and
a wrong guess leaves mutants unrun. registry.py holds exactly one such class:
Tap, 30 lines of 783, which carries no recorded mutants at all —
its 24 measured symbols sum to 7,309,375 against a file total of 7,309,374. So a class that
contributes nothing to the cost blocks the split of everything that does. Eleven files in
boost_cli/core are in the same position, 18.6% of all mutation work.
The fix is proof, not a guess. cmd_weights writes
millis_by_symbol[file] only when every mutant of the file was timed, and
it derives each key by the exact inverse of the mangling pattern_for applies. So
a complete record both covers every mutant and addresses it. measured_partition
lifts the AST refusal for a module a real run has cleared, under one guard: every recorded
name must be a top-level def, since a name from inside a class would round-trip
to a pattern matching nothing. Over both committed weight generations and all eleven
class-bearing files, no recorded name has ever come from inside a class. The split
still runs over the AST rather than the record, so a function added since the measurement
keeps its unit; and cmd_merge fails closed on any unrun mutant, so the worst
case is a loud red build rather than a score computed over a short set.
Measured result. On #1022's weights the pack goes from TOO TIGHT at every shard count
to split files : registry.py, store.py and a 58.8-minute tail at ten shards,
with catalog.py (5,914,538 ms) as the new floor. On the weights committed today
nothing changes at all — registry.py is under the even share, so it is not
split and the plan is byte-identical.
Developer experience & maintainability
// planned · free tooling to catch issues earlier & keep the code legiblePromote nav / footer into the shared style system
Write-up · promote-nav-footer-into-the-shared-style-system.md
Shift-left gate — pre-commit + pre-commit.ci
Layering guard — import-linter
Typo detection — codespell
Docstring coverage — interrogate
Performance-regression gate — pytest-benchmark
Modernization smells — refurb + pyupgrade
Coverage dashboard — Codecov (free for OSS)
Pin the lint toolchain so a release can't redden the gate
Runner egress monitoring — StepSecurity Harden-Runner
Write-up · runner-egress-monitoring-stepsecurity-harden-run.md
Required-status-checks list is prose-only, and already stale
No issue/PR templates or a code of conduct
make boost mcp the whole setup, and put all three kinds behind it
New BDD suite has zero CI wiring
A project-scoped install registers its MCP servers machine-wide
boost bmad needed Node and a per-project install before it did anything
boost adapt — render a skill as another framework's agent source
boost adapt --to langgraph — third framework renderer
boost adapt — multi-agent skills → crews/graphs, not one Agent
boost run — search → adapt → a live agent doing the task, in one command
MCP — make agents search boost before reinventing a skill
boost install resolves a skill's requires: closure
MCP-aware skills — declare and wire an .mcp.json on install
roadmap.html goes stale on every rebase, so a card and a merge race redden the whole matrix
make lint reports success when actionlint fails — and says it wasn't installed
MCP — check for a skill when a task starts, not only when authoring one
The second silent skip — actionlint runs, and checks no run: block at all
boost chat cites its sources by a number it never prints
Write-up · chat-cites-sources-by-a-number-it-never-prints.md
MCP — one benefit, one observable trigger (and stop routing through boost_info)
MCP — answer the veto that overruled the trigger ("a skill already matched")
boost-first — the one rule boost authors, offered opt-in at boost mcp register
The MCP instructions understated what a search costs by 100x
install dead-ends on a registry that vendors its own skills
boost completions completes command names and nothing else
install refused an ambiguous name and offered no way to answer it
boost discover <query> asks GitHub, instead of filtering whatever boost index happened to sample
Share the catalogue instead of making everyone re-tap it
boost-first carried the trigger that had already fired and lost — and could never be updated
boost serve becomes a searchable, faceted catalogue with a graph of the taps
completions --install could delete the config between its own markers
browse could not search for two words, and the fix reshaped the whole browser
One design system across search and browse
The smart rerank pays the LLM again for a search it already answered
The box drew 108 columns into an 80-column pane, and --help never asked how wide the pane was
The hints still run past the pane, and the worst one is pinned by six test files
MCP has no way to read a skill before installing it
clean counts failed removals as cleaned, journals the inflated count, and exits 0
Write-up · audit-clean-counts-failed-removals-as-cleaned-journals-the-inflate.md
config/policy set store type-unchecked values; consumers crash exit 70 and pin_only no freezes installs
Write-up · audit-config-policy-set-store-type-unchecked-values-consumers-cras.md
Corrupt settings/config/state JSON silently read as empty, then clobbered on the next write
Write-up · audit-corrupt-settings-config-state-json-silently-read-as-empty-th.md
Five exit-70 crashes on bad paths/data: catalog --export, serve --port, count, replay, infer -o
Write-up · audit-five-exit-70-crashes-on-bad-paths-data-catalog-export-serve.md
distill's heuristic merge drops repeated ``` fences/braces, writing a structurally corrupt SKILL.md
Write-up · audit-distill-s-heuristic-merge-drops-repeated-fences-braces-writi.md
create/distill/infer/absorb --install silently replace an installed (unpinned) skill and flip its lock provenance to local
Write-up · audit-create-distill-infer-absorb-install-silently-replace-an-inst.md
hooks remove -n cannot find hooks boost itself added (unknown events skipped, embedded # boost: mangles the name)
Write-up · audit-hooks-remove-n-cannot-find-hooks-boost-itself-added-unknown.md
install --dry-run promises agents the real install never writes (antigravity-cli copy, antigravity materialize) and omits the MCP plan
Write-up · audit-install-dry-run-promises-agents-the-real-install-never-write.md
Rules/workflows reported as 'skills': uninstall claims a store dir that never existed
Write-up · audit-rules-workflows-reported-as-skills-uninstall-claims-a-store.md
install --local writes the repo's .mcp.json silently — the 'recorded N servers' report never runs
Write-up · audit-install-local-writes-empties-the-repo-s-mcp-json-silently-of.md
mcp register's boost-first consent names one file but writes every agent
Write-up · audit-mcp-register-s-boost-first-consent-names-only-the-registered.md
catalog.resolve_one's duplicate-name hint advertises --path to commands that reject it
Write-up · audit-catalog-resolve-one-s-duplicate-name-hint-advertises-path-to.md
adapt/run/stats/edit/tag/export reject the tap:name qualifier that info/install accept and adapt's own hint recommends
Write-up · audit-adapt-run-stats-edit-tag-export-reject-the-tap-name-qualifie.md
tap --dry-run is silently ignored outside --catalog: SPEC and --defaults clone for real
Write-up · audit-tap-dry-run-is-silently-ignored-outside-catalog-spec-and-def.md
out.warn defaults to stdout: infer/absorb corrupt > SKILL.md; search/explain/context warnings pollute piped stdout
Write-up · audit-out-warn-defaults-to-stdout-infer-absorb-corrupt-skill-md-se.md
AI degrade note blames PATH/API keys regardless of cause; several commands fall back with no note at all
Write-up · audit-ai-degrade-note-blames-path-api-keys-regardless-of-cause-sev.md
Declined confirms never name -y/BOOST_ASSUME_YES; snapshot, clean, infer/distill/absorb and sync reject --yes
Write-up · audit-declined-confirms-never-name-y-boost-assume-yes-snapshot-cle.md
Dry-runs disagree with the real run: compact, heal and onboard previews mispredict
Write-up · audit-dry-runs-disagree-with-the-real-run-compact-counts-bytes-it.md
Name slugging is inconsistent: distill -o accepts what import rejects; create/profile slug silently
Write-up · audit-skill-profile-name-slugging-is-inconsistent-distill-o-accept.md
git ops never set GIT_TERMINAL_PROMPT=0, so a 404/private repo prompts for credentials
Write-up · audit-git-operations-never-set-git-terminal-prompt-0-so-a-404-priv.md
cli.py COMMANDS summaries and parser help contradict behavior across ~11 commands
Write-up · audit-cli-py-commands-summaries-and-parser-help-contradict-behavio.md
info/stats/explain render a smaller shape for rules and workflows than for skills
Write-up · audit-info-stats-explain-render-a-different-smaller-shape-for-rule.md
taps/outdated/decay/policy/snapshot --json emit display strings as machine fields
Write-up · audit-taps-outdated-decay-policy-snapshot-json-emit-display-string.md
--json accepted but ignored: cohort/config/policy set, focus, profile, replay rollback, who empty state
Write-up · audit-json-accepted-but-ignored-on-many-branches-cohort-config-pol.md
Missing --json on doctor, test, health, changelog, trending, log, hooks list and friends; bundle install lacks --dry-run
Write-up · audit-missing-json-on-doctor-test-health-changelog-trending-log-ho.md
Name-miss errors: unknown tap qualifier never named, no close-match hint, tap tokens pollute suggestions
Write-up · audit-name-miss-errors-unknown-tap-qualifier-unnamed-no-close-matc.md
Six count flags (tap/chat/absorb/lint/changelog/hooks) accept 0 and negatives
Write-up · audit-six-count-flags-tap-chat-absorb-lint-changelog-hooks-skip-ut.md
Project scope seams: uninstall/verify/list/info disagreed with what install --local wrote
Write-up · audit-project-scope-seams-uninstall-verify-list-info-reinstall-dis.md
Project scope in an unmarked tree does not walk up, so src/ becomes its own project
scopes.resolve_base walks up for a VCS marker and, finding none, falls back to the
directory it started in · without walking up. So in a project that is not a git/hg/svn
checkout, every subdirectory is its own project: install brainstorming --local run from
proj/src writes proj/src/.boost/skill-lock.json and proj/src/.claude/skills/,
even when proj/.boost/skill-lock.json already exists; and list --local,
info, doctor, verify and bare uninstall run from
proj/src cannot see what was installed at proj. mkdir .git fixes
both, which is what makes it easy to miss.
This is not the reader/writer split that fix(scope) closed · readers and writers now
agree, and that is the point: they agree on the cwd rather than on the project.
project_root's docstring says walking up is the whole point ("install --local
run from src/deep/nested must write into the repo's .claude/skills, not create
a stray one three levels down"), and that promise is simply not kept once the marker is absent.
The fix is a choice, not a bug fix, which is why it is its own card. Option (a): make the fallback
walk up for an existing .boost/skill-lock.json · but a lock-file marker cannot
help the very first install, which is the one that creates it. Option (b): let the fallback walk up to
the nearest ancestor that is not $HOME · too greedy, it would swallow unrelated
sibling directories. Option (c): keep today's behaviour and say so · have
install --local warn once when it is about to create a project in a directory with no VCS
marker, naming the directory, so the second lock is a decision rather than a surprise. (c) is the
cheapest and the most honest; (a) is worth pairing with it for every command after the first.
Found by the adversarial verification of PR #998, which reproduced it end to end. $HOME
is still never a project and deletion is still gated by resolve_in_base, so this is a
wrong-directory bug, not a safety one.
A rule or workflow installed --local reads as a user one everywhere but list --local
Rules and workflows installed with --local materialize into the repo but are recorded
in the user lock, tagged scope: project and base: <repo>,
because a project lock holds skills and nothing else
(store._check_scope_conflict). boost list --local now filters them with
scopes.owned_by, and it is the only reader that does.
lockfile.installed_rules() and installed_workflows() have a dozen other
consumers — store, complete, taps,
team, safety, quality, pkg — and none of
them ask whose repo a row belongs to. So boost doctor and boost verify
count another checkout's rule as this machine's, drift and health grade
it, tab-completion offers it, and plain boost list prints it under
“installed rules” with nothing saying it belongs to a directory the user may not be
standing in. The failure is quiet in both directions: a row for a repo that has since been
deleted never goes away, and a row for the repo you are in looks identical to one for a
repo you are not.
Worth settling the design before the sweep, because the honest answer differs per command. A
scope column on plain list's rule and workflow tables is cheap and makes
the ambiguity visible (boost info already prints scope project).
Whether verify and doctor should grade a rule belonging to
another repo is a separate question — they cannot read its materializations from here, so
counting it is arguably the bug and skipping it with a note the honest fix. Split out of
audit-project-scope-seams-uninstall-verify-list-info-reinstall-dis, where the
evidence was gathered.
export -o, cohort create and profile save silently overwrite existing outputs and still say created/saved
Write-up · audit-export-o-cohort-create-and-profile-save-silently-overwrite-e.md
Stray positionals and inapplicable flags silently ignored across import, config, policy, trust, log, schedule, hooks, snapshot
Write-up · audit-stray-positionals-and-inapplicable-flags-silently-ignored-ac.md
out.table clips data columns to an assumed 80 columns when stdout is a pipe; narrow TTYs clip IDs/hashes
Write-up · audit-out-table-clips-data-columns-to-an-assumed-80-columns-when-s.md
output._wrap_tokens splits punctuation off backtick spans, inserting a stray space
Write-up · audit-output-wrap-tokens-splits-punctuation-off-backtick-spans-ins.md
Sweep: positionals/--json lack help strings and no command help shows examples (~30 cmds)
Write-up · audit-sweep-positionals-json-lack-help-strings-and-no-command-help.md
Singular/plural misses (“1 skills”, “1 issue need”, “1 skill pass”) across six commands
Write-up · audit-singular-plural-misses-1-skills-1-issue-need-1-skill-pass-ac.md
log/simulate/changelog headings echo the raw qualified argument, printing the tap twice
Write-up · audit-log-simulate-changelog-headings-echo-the-raw-qualified-argum.md
Stale prose after shipped changes: catalog --export, live discover, 464 count, Apache-2.0
Write-up · audit-stale-prose-after-shipped-changes-catalog-export-live-discov.md
boost ROOT: CLI audit findings (2026-08)
boost absorb: CLI audit findings (2026-08)
boost adapt: CLI audit findings (2026-08)
A subagent named after a declared tool renders modules that compile but cannot run. With a
subagent grep and tool Grep: crewai emits @tool("grep") def
grep then grep = Agent(...), so reviewer_1 = Agent(tools=[read,
grep]) hands the Agent where a tool belongs (stub run: TypeError); langgraph assigns
grep = create_react_agent(...) inside build_mycrew, making it local, so the
earlier tools=[read, grep] raises UnboundLocalError. _unique_idents
(core/adapters.py:188-202) dedups only among agent specs and _unique_tools
(:205-212) allocates stub names independently. Fix: pass the tool ident set into
_unique_idents as pre-reserved names (or prefix stubs tool_<name>) in
render_crew/render_graph, plus a golden test that executes a colliding
render against stubs.
docs/commands.html brackets required options as optional. Line 369 shows
boost adapt [--to FRAMEWORK] … while adapt --help prints an
unbracketed --to FRAMEWORK and omitting it exits 2 — and the verify pass found it
is broader than adapt: evolve's required --feedback and catalog's required
mutually-exclusive group render all-optional too. Pure generator bug:
scripts/build_command_reference.py:122-126 brackets every option unconditionally. Emit
required options unbracketed (a required group as (--a | --b)), prefer the short flag
like argparse, then make generate — the --check gate holds it after
that.
Colon-form model ids get double-prefixed for the LiteLLM targets.
--model anthropic:claude-x — the form langgraph accepts and emits —
renders llm=LLM(model="anthropic/anthropic:claude-x") for crewai and the same for
agents-sdk; multi-agent crews inherit it via adapters.py:376.
_litellm_model's docstring says a provider-qualified value passes through, but the code
checks only /. Fix: replace the first : with / before deciding
to prefix (mirror of _langchain_model); document accepted syntaxes in
docs/adapters.html's --model paragraph (~line 344, slash form only today),
and regenerate docs/commands.html only if the argparse help changes.
adapt -o and run --print -o write generated source mode
0600 — and re-rendering over an existing 0644 file silently downgrades it, unlike a shell
redirect (-rw------- vs -rw-r--r-- under umask 022).
util.atomic_write_text (core/util.py:91-116) inherits mkstemp's 0600, right
for the lock/config it was written for, wrong for source the user asked boost to write. Fix: add an
optional mode parameter (fchmod the temp fd before os.replace), keep 0600
the default, and have cmd_adapt (pkg.py:1753-1761) and cmd_run
(run.py:62) pass the umask default. Still open — an implementation of exactly this
shape (path-based os.chmod, not os.fchmod, learned the hard way: the latter
raises on Windows) shipped and then was reverted from PR #728 after three rounds of Windows-only
windows-latest CI failures in the "unit + functional (90% coverage gate)" step that
neither pytest-cov nor a temporary diagnostic artifact-upload commit could surface a cause for —
the session driving that PR could not read Windows job logs (capped and consumed by
harden-runner's own diagnostic noise) or download the diagnostic artifact
(productionresultssa*.blob.core.windows.net blocked by that session's network egress
policy) to see the actual failure. The other three findings landed clean on every platform. Whoever
picks this back up needs either a session with working Windows CI log access, or to reproduce
locally on a real Windows box.
Found by the 2026-08 CLI audit (clusters adapt-ident-collision,
docs-required-flag-synopsis, adapt-model-id-syntax,
generated-file-mode); repro in the audit log.
boost bmad: CLI audit findings (2026-08)
boost browse: CLI audit findings (2026-08)
boost bundle: CLI audit findings (2026-08)
boost catalog: CLI audit findings (2026-08)
boost changelog: CLI audit findings (2026-08)
boost chat: CLI audit findings (2026-08)
boost cohort: CLI audit findings (2026-08)
boost completions: CLI audit findings (2026-08)
boost config: CLI audit findings (2026-08)
boost context: CLI audit findings (2026-08)
boost deps: CLI audit findings (2026-08)
boost discover: CLI audit findings (2026-08)
boost edit: CLI audit findings (2026-08)
boost evolve: CLI audit findings (2026-08)
boost explain: CLI audit findings (2026-08)
boost export: CLI audit findings (2026-08)
boost import: CLI audit findings (2026-08)
boost index: CLI audit findings (2026-08)
boost install: CLI audit findings (2026-08)
Rule/workflow lock entries are name-keyed across scopes (med). With the
benchmarking rule at user scope, install benchmarking --local in a project fails
“Error: benchmarking is already installed / hint: boost reinstall benchmarking
to force” — the project has no copy, and the hint would reinstall the user one. The
reverse direction blocks too, and skills coexist fine (separate project lock). Worse, the
--force escape overwrites the user-scope lock entry with the project one, orphaning the
user materializations so uninstall can no longer clean them. _install_rule
(store.py:836-839) and _install_workflow (store.py:1074) gate
on a name-only lookup with no scope/base comparison. Fix: key entries by scope (or compare
existing scope/base before raising), word the error “already installed at user
scope”, and refuse a cross-scope --force overwrite without cleanup. Docs:
README's install-scope section (~301-328) and
docs/roadmap/items/install-scope-user-or-project.md.
(Cluster cross-scope-name-block.)
--path says “under path” but matches suffix-only
(low). --path plugins/tdd/skills is refused while the error's own hint lists
plugins/tdd/skills/test-driven-development — a path that is under it.
Suffix matching is the shipped design (install-path-disambiguation, PR 483); the wording
is the defect. Reword the raise in catalog.py:~502 to “no copy of X whose path ends
with Y” and hint “pass a trailing segment of one of: …”.
(Cluster install-path-prefix-match.)
The MCP offer never shows the runnable command (low). The server row prints
only demo-echo npx though the sidecar declares npx -y
@example/demo-echo-mcp plus env, and on decline the hint is a literal elided
claude mcp add …; the full argv only prints when the host CLI is
missing. _offer_mcp renders how from spec['command'] alone
(pkg.py:161-164) and mcpdecl.register_argv already exists
(pkg.py:201) — render command+args, print the joined argv on decline, and indent
the confirm prompt to match its neighbours. (Cluster mcp-offer-command-detail.)
The typosquat warning prints three times (low).
install NeoLabHQ/context-engineering-kit:test-driven-development --dry-run prints the
identical “closely resembles test-driven-development
(sickn33/antigravity-awesome-skills)” warning 3×, one per mirror copy in the
look-alike tap. De-duplicate find_confusions on (name.lower(), tap)
(typosquat.py:79-87) so the [:3] slice in _warn_confusions
covers three distinct look-alikes. Found by the 2026-08 CLI audit (cluster
typosquat-warning-dupes); repro in the audit log.
Status (2026-09). Three of the four clusters shipped as described above:
install-path-prefix-match (catalog.py wording), mcp-offer-command-detail
(mcpdecl.command_line renders the full command+args, the decline path prints the real
argv), and typosquat-warning-dupes (find_confusions dedupes on
(name.lower(), tap)). cross-scope-name-block got the narrower of the
fix's own two options: store._check_scope_conflict now refuses a rule/workflow install
whose name collides with an existing lock entry recorded under a different scope/base —
naming the real location ("already installed at user scope") and refusing even under
--force, which closes the silent-corruption half of the bug (a forced cross-scope
install used to overwrite the other scope's lock entry, orphaning its materializations). What is
still missing is the other half: rules and workflows still cannot coexist across scopes the
way skills do, because they share one lock keyed by bare name with no per-location table — skills
got a separate projectlock.py when project scope was added, rules/workflows never did.
Giving them the same treatment (a rules/workflows section in
projectlock.py, wiring _install_rule/_install_workflow and
their uninstall/sync counterparts through it for project scope) is real coexistence but is its own,
larger change, and belongs in its own card rather than folded into a bugfix PR.
boost log: CLI audit findings (2026-08)
boost mcp: CLI audit findings (2026-08)
boost onboard: CLI audit findings (2026-08)
boost outdated: CLI audit findings (2026-08)
boost preview: CLI audit findings (2026-08)
boost profile use: CLI audit findings (2026-08)
boost protocol: CLI audit findings (2026-08)
boost pulse: CLI audit findings (2026-08)
boost quickstart: CLI audit findings (2026-08)
boost recommend: CLI audit findings (2026-08)
boost replay: CLI audit findings (2026-08)
boost run: CLI audit findings (2026-08)
boost schedule: CLI audit findings (2026-08)
boost search: CLI audit findings (2026-08)
boost simulate: CLI audit findings (2026-08)
boost sync: CLI audit findings (2026-08)
boost tag: CLI audit findings (2026-08)
boost tag swallows unknown flags and misreads them as operands. tag brainstorming --verbose prints the current tags and exits 0 — the flag is consumed as a removal of the tag -verbose; tag --verbose gives "Error: --verbose is not installed" (the flag becomes a skill name); verification found a third hole: tag brainstorming --list silently discards the skill-name operand and lists all tags. Cause: cmd_tag's manual split (boost_cli/commands/info.py:988-993) whitelists only --list/--json/-h/--help; every other --x token falls through as an operand. Every sibling command rejects unknown options with "unrecognized arguments" exit 2.
And the mutation path has no before/after check. tag brainstorming -nosuch removes a tag that was never present — silent, exit 0; tag brainstorming +x -x prints ✓ and writes the lock plus a journal event for a net no-op (changed is set per-token at info.py:1027-1041, never compared to the before set); "+with space" is accepted as #with space; +Design and #design coexist. The shipped roadmap item robust-tag-argument-parsing (PR 94) built this manual split — these are residual holes in it, not a duplicate.
Fix in cmd_tag: hand any token starting with -- (or -letter that is not a tag operand) to argparse so it errors; compute changed = sorted(tags) != sorted(before); print a one-line notice for removing an absent tag; reject whitespace in tags; document or fold case; error when a name is given with --list. Regenerate docs/commands.html if the help text gains the tag grammar.
Found by the 2026-08 CLI audit (cluster tag-arg-parsing); repro in the audit log.
Partly landed — PR 735. The correctness half shipped: any unrecognized
-- token now reaches argparse (unrecognized arguments, exit 2),
--list with a skill name is a named error, whitespace in a tag is rejected, and
changed is a before/after set comparison in the new
lockfile.apply_tag_mods, so +x -x no longer writes the lock and a
journal event for a net no-op. Still open, and why this card stays
inflight: the one-line notice when -tag removes a tag that was
never present (the remove branch is still a silent no-op), and documenting or folding tag
case (+Design and #design still coexist). Both are UX asks rather
than correctness bugs, which is why the PR left them.
boost tap: CLI audit findings (2026-08)
boost taps: CLI audit findings (2026-08)
boost unpin: CLI audit findings (2026-08)
boost untap: CLI audit findings (2026-08)
boost update: CLI audit findings (2026-08)
boost who: CLI audit findings (2026-08)
August 2026 full-CLI audit: every one of the 81 boost commands exercised and verified
Per-item categories in search/browse/info — not just a ★ curated bool
From a user request: “proper categories for skills (can't have all of them listed as just
curated)”. They are right about the item level: the only taxonomy a catalog entry carries is
a boolean. A boost search row shows name, kind, tap, description and at most a
★; boost recommend's no-match fallback is literally headed
“curated picks”; boost info prints no category at all. Across a
real install of tens of thousands of items, “starred or not” is the entire
classification a user can see or filter by.
What the code confirms. catalog._make_entry stamps "curated": curated
onto every entry (boost_cli/core/catalog.py:119, signature at 105–106) — and that
bool is per-tap, from Tap.curated (core/registry.py:23), set by
tap --defaults or by anyone passing --curated
(commands/taps.py:126) — a trust star, not a classification. Category-like data does
exist, but only per tap: data/registries.json rows carry one (487 registries, 21
values; general alone covers 127), and exactly two surfaces read it —
browse's row badge via _tap_categories
(commands/discovery.py:936–941, whose own docstring says “catalog
entries themselves carry no category, only their tap does”; badge appended last in
_row_badges, discovery.py:961–963, so narrow panes drop it first, and taps
outside the bundled 487 get none) — and boost serve's web facets
(core/serve.py:65). cmd_search renders only the star
(discovery.py:179) and takes no filter flag; info shows frontmatter tags when present
(commands/info.py, the meta.get("tags") kv) but no category, and its
--json has no such field. An item's own frontmatter category/tags
ride along invisibly in entry["meta"] and the substring search_blob
(catalog.py:131, 621–627), so they can match a query yet can never be displayed or filtered.
Proposed fix. Stamp a first-class category on each entry at scan time in
_make_entry (catalog.py:105–132): the item's frontmatter category
(or first tag) when declared, else inherited from its tap's registry category — and bump
catalog.CACHE_FORMAT so hundreds of existing tap caches backfill without a re-tap, per
the versioned-cache rule. Then surface it where a category would live: a badge in
search rows and a --category filter on
search/browse/recommend, a kv row plus JSON field in
info, and browse's existing badge switched from tap-level to the entry
field (which also gives un-bundled taps' items a label for the first time). ★ keeps meaning
curation/trust only. Consumers must degrade cleanly when category is absent (old
caches, synthesised entries), same as the content digest rule.
Docs: regenerate docs/commands.html for the new flags; no other doc names categories.
Found by the 2026-08 CLI audit (cluster catalog-categories-beyond-curated, filed from
the user's request); repro in the audit log. Verified against source 2026-08-31.
Status (2026-09-01). Landed: the category stamp at scan time
(catalog._entry_category, own frontmatter category → first
tags entry → tap's registry category), CACHE_FORMAT bumped to 2 so
existing caches backfill on next scan, a --category filter on
search/browse/recommend
(catalog.matches_category/filter_by_category), info's kv row
and --json field, and browse's row badge switched from the tap-level
lookup to the entry's own field (falling back to the tap lookup for a cache not yet rescanned).
Not done: the badge in plain boost search rows. That row's column widths
(out.search_layout/format_search_row) are a tuned, heavily-pinned budget
system (drop order, per-cap name shrinking, a reserved curated tail) — working it out safely needs
its own pass rather than a bolt-on inside this PR. Left inflight rather than
shipped for that reason; the next claim on this item is scoped to exactly that piece.
The last command that blamed boost for the user's typo
One incidental keyword decides the boost bmad track
The BMAD router reads each prompt alone, so a pasted log, an “ok update both and rerun” and a yes/no question all get a banner
boost bmad gives every track the build contract, so a review is told to finish a change
bmad route never checks autopilot state, so any hook off missed keeps routing
The autopilot routes docs at a skill BMAD 6.12 no longer installs, and no test would notice
On Gemini, the bmad on briefing reaches the user and never the model
Persona descriptions say Use PROACTIVELY, so the router's silence governs only the banner
out.err() judges colour by stdout while writing to stderr
boost quickstart: a rerun could move already-tapped registries to their vectors
A rerun of quickstart still leaves an already-tapped registry at the commit it was tapped at. registry.add_many skips a configured tap (already tapped), so a registry first tapped before quickstart pinned anything, or overtaken by a weekly republish, stays where it is, and shards.sync then refuses its shard (tap is at X, shard is for Y). The audit-quickstart-findings fix flags that refusal and names boost update --shards, which already moves such taps through shards.ingest (download and verify first, then move the tap). That was the smaller half of the card; its other half was to have quickstart do the move itself. Fix: on a rerun, hand the configured taps the manifest publishes to shards.ingest instead of shards.sync, so the vectors land in one command. Decide first whether quickstart may move a tap the user tapped on purpose at another commit: boost update treats a pin as a promise, and a tap pinned with boost tap --at should keep it. Pin that choice in tests/functional/test_cli_quickstart.py beside the commit_moved tests.
uninstall --local cannot remove a rule or workflow that list --local shows
Write-up · uninstall-local-cannot-remove-a-rule-or-workflow-list-local-shows.md
Docs-site & content quality
// planned · free tooling for the Pages site, README & proseMake the roadmaps discoverable
Lighthouse CI on the Pages site
Accessibility audit — pa11y-ci / axe-core
HTML validation — html-validate
Broken-link & anchor checking — lychee
The OpenSSF badge playbook
Prose & terminology linting — vale
Markdown consistency — markdownlint-cli2
Theme-asset linting — stylelint + eslint
Post-deploy smoke — headless load check
Command reference documentation site
Surface every docs/*.html page from the main page
The engine had no architecture diagram — and the one written rule was documented backwards
Fix overflowing node text in the RAG diagram (mcp-hub.html)
Docsite audit — stale counts, dev-noise footers, and a nav that breaks on mobile
A GIF carousel touring one flagship command per group
demo.yml has failed every run since it landed — vhs-action cannot install ffmpeg
No install doc ever said how to upgrade, so users guessed install --upgrade — which no-ops
Nothing tells a user semantic search is off — not the README, not search, not /mcp
roadmap.html grew 36% in one session and nothing bounds it
The Lighthouse budget passes on noise, not on margin
an explainer page for the LangChain / LangGraph / LangSmith integration
The performance gate was measuring a page nobody is served
Expanded card bodies overflow the roadmap board sideways
The performance gate flips on byte-identical input
Compatibility & install integrity
// planned · free tooling to prove boost installs & runs everywhere it claimsWindows in the CI matrix
Clean-env install smoke — pip & pipx
Lowest-version resolution — uv --resolution lowest-direct
Write-up · lowest-version-resolution-uv-resolution-lowest-d.md
Pre-release Python canary — 3.14t free-threaded
Startup & import-time budget — -X importtime
Package-metadata validation — twine check + friends
Write-up · package-metadata-validation-twine-check-friends.md
Hash-pinned, reproducible toolchain — requirements/*.txt
What Gemini actually receives from boost, audited
One command, every env — nox
boost hooks learns a second host — and finds two bugs upstream
Harden boost mcp launch against macOS Obj-C fork aborts
Self-harden every boost process against the macOS fork-safety abort
self-update is non-functional for pip/pipx installs
self-update said "already up to date" without asking PyPI
sync reported success for a link it had just refused
Two crashes that should have been messages
Order the server name before -e flags in `boost mcp register`
Gemini CLI as a first-class agent target (skills, rules, workflows, MCP)
production-ready LangChain / LangGraph / LangSmith integration
ship the LangChain integration inside the wheel, behind a [langchain] extra
sanitize agent frontmatter for Gemini instead of copying it verbatim
BoostRetriever advertised a source that does not open, and k=0 returned nothing forever
garrytan/gstack — tap it first, then learn to coexist with it
73 reads and writes left the text encoding to the locale, and the Windows jobs were the only thing that noticed
917 text writes leave the line ending to the platform, and only some of them are bugs
Text mode decides two things, and naming the encoding fixes one of them. It also translates
\n to the platform separator on write, so a file written with
encoding="utf-8" on a GitHub Windows runner still holds CRLF. The sibling card
text-io-leaves-its-encoding-to-the-locale closed the encoding half across 414 files; this
is the other half, and it is deliberately not the same kind of sweep.
#1012 hit both, one after the other, and the first fix hid the second. Naming the encoding
turned the three tests (windows-latest, 3.1x) jobs green on the assertion that had failed
— and the next run put them straight back to red on the line endings of the same line. Two
platform defects in one call, and no way to see the second until the first was gone.
The count is 917, and a blanket pass would be wrong. 858 in tests, 32 in
boost_cli, 21 in scripts, 6 in evals. Most of those writes
produce files whose line endings nothing ever compares, so pinning
newline="\n" on all of them is 917 edits of which the overwhelming majority are churn
and none is a fix — and churn in a diff is what makes the real change unreviewable. Readers need
nothing at all: universal newlines already fold CRLF to LF on the way in, which is why the guard that
ships with the encoding card walks writers only.
So the work is the scoping, not the edit. The subset that matters is writers whose output is
read back and compared — against a string, a hash, a golden file, or the same report printed to a
log beside it. Three shapes are known to qualify and are a good place to start the inventory:
a file written here and read back by an assertion in the same test; anything hashed or diffed
(fingerprint, the generated-file --check gates, golden fixtures); and
anything a second process parses line by line, which is how
$GITHUB_OUTPUT got onto the list. scripts/mutation_shards.py is already
pinned and already guarded by TestTheScriptPinsItsNewlines — extending that guard a
scoped set at a time, with the reason each set is in it, is the shape this should take.
What would make it falsifiable. A Windows job that asserts byte equality, not string
equality: today a round trip on Linux passes whether the newline is pinned or not, which is exactly
the class of test the encoding card's own verification pass caught twice. The AST walker
(unpinned_newline in tests/unit/test_text_io_names_its_encoding.py) is the
platform-independent half and already exists; what it cannot tell you is which of the 917 are bugs.
Skill-content trust & safety
// planned · boost's core threat model — the third-party skills it installs run inside an agentPrompt-injection scanning of skill Markdown
Integrity verification — boost verify
The MCP boost_install tool skipped the injection scan the CLI runs
Crash reports carried API keys in cleartext
Tap signing & provenance — Sigstore / minisign
Runtime hallucination guardrail for boost explain
Typosquat & name-confusion detection
Update-diff before apply
Project scope — refuse to write through an escaping symlink
Secret & PII scanning of installed skills
Lockfile enforcement & commit pinning
Capability manifest & least-privilege policy
Write-up · capability-manifest-and-least-privilege-policy.md
install_from_path bypasses pin & policy checks
Write-up · install-from-path-bypasses-policy-and-pin-checks.md
Path traversal via unsanitized rule/workflow name
A CodeQL job rename silently blocked every merge
boost audit --skills — a trust/staleness report for installed skills
sbom.yml has never run — it waits for an event GITHUB_TOKEN cannot emit
The main ruleset is inert — its ref pattern is refs/heads/"main", quotes included
The release trigger was reachable from a fork — branches: filters head_branch, not the event
The code_scanning ruleset rule can go back on — but only scoped to CodeQL
The SBOM can declare a different version than the release it is attached to
rules and workflows install, then cannot be governed
boost serve echoed the request path back into its 404 body
denied_capabilities policy never applied to rule/workflow installs
trust verify labels a manifest tampered after signing by a TRUSTED key 'untrusted'; sweep exits 0
Write-up · audit-trust-verify-labels-a-manifest-tampered-after-signing-by-a-t.md
util.rmtree's retry hook chmods a symlink's target, outside the tree it is deleting
util.rmtree installs an error handler so a read-only file cannot strand a delete: on
failure it chmods the path and retries. The handler is handed the path that failed, and
when that path is a symlink the chmod follows it · so deleting a
directory that happens to contain a link to a file elsewhere changes that file's permissions.
Measured, because the first draft of this card guessed at the shape and guessed wrong. The hook only
fires once a delete has already failed, and unlinking a symlink does not fail while its
parent directory is writable · so in the ordinary case nothing happens at all: the link goes,
the target keeps its mode byte-for-byte. Make the parent read-only (0o500) and the
unlink fails, the hook chmods the link, the target outside the tree goes
0o400 → 0o200, the retry fails again and rmtree raises.
The link is still there. So the only lasting effect of the hook on this path is the permission
change on a file it was never asked to touch · the earlier claim that “the delete
itself correctly removes only the link” describes a run in which the hook never fires.
The containment guards do their job: scopes.resolve_in_base and contains
still refuse to remove anything outside the base, and the verification of PR #998 confirmed a planted
victim file survives every escape attempt byte-for-byte. What is not guarded is the
permission change on the way past, which happens before containment is ever consulted
because it is inside the retry hook.
Pre-existing and untouched by the project-scope work, but newly easier to reach: bare
uninstall can now act in an unmarked directory, so a repo copy that carries a symlink
into the user's home is one command away. The fix is small · use
os.chmod(..., follow_symlinks=False) where the platform supports it, and otherwise skip
the retry entirely for a path that os.path.islink reports, since a symlink's own mode is
not what blocked the unlink. A test wants a link inside the tree pointing at a
0o400 file outside it, asserting the target's mode is unchanged after the delete.
Implementing it turned up a second, worse instance of the same bug, which this card did not
know about. Hand util.rmtree a symlink as its argument and
shutil.rmtree refuses it by calling the error hook with func set to
os.path.islink. The old hook chmodded straight through the link and then called
os.path.islink(path), which answers True without raising · so the hook returned,
rmtree returned, and the caller was told a tree had been removed when nothing had.
Measured: link intact, target directory intact with its contents, target's mode 0o755 → 0o200.
Silent, and independent of any 0o500 — the first variant at least raises.
The repair the card proposed was the wrong one.
os.chmod(..., follow_symlinks=False) needs lchmod, which
os.supports_follow_symlinks does not report everywhere this runs (measured True on
darwin), so it buys a platform branch that cannot be exercised on the runner — an unkillable
mutant by construction. It is also treating a mode that was never the blocker: unlinking is gated by
the parent directory's write bit, not the link's own mode, which is why the retry fails a second
time in the first variant. Re-raising the exception the hook was handed is correct on both counts
and needs no branch.
boost attest: CLI audit findings (2026-08)
boost trust: CLI audit findings (2026-08)
uninstall --local deletes any in-repo directory the committed lock names
Write-up · project-uninstall-deletes-any-in-repo-directory-the-lock-names.md
The add_many concurrency probe measured a race, not the pool
test_it_actually_runs_concurrently counted the peak number of fake clones in flight and
asserted it exceeded one. The fake clone is one mkdir and a small write, so the window
in which two of them overlap is microseconds wide · on a loaded runner each worker finished
before the next was scheduled, the peak stayed at 1, and the required tests (macos-latest,
3.14) check failed with assert 1 > 1 against a ThreadPoolExecutor
that was doing exactly its job. Observed on PR #998, where it blocked a merge for a defect that did
not exist.
The failure said the wrong thing, which is the part that matters. "This is serial" and "this
was too fast to catch overlapping" are different diagnoses, and the second wastes whoever reads it.
A required check that can fail for a reason unrelated to the change also trains the next person to
re-run rather than read · which is how a real regression gets waved through.
A threading.Barrier(jobs) inverts the dependency. Every worker blocks until
jobs of them have arrived, so the assertion stops being about timing: a pool of four
trips the barrier and the call returns, and a pool narrower than four cannot — the
first worker waits for peers that will never come and the barrier times out. Slower hardware makes
the test more reliable rather than less, which is the opposite of the property it had.
Paired with test_the_concurrency_probe_fails_when_the_pool_is_serial, which drives
add_many at jobs=1 so the probe's own ability to fail is itself asserted,
and verified by hand against a max_workers=1 executor: it fails with "the pool never
had 4 clones in flight".
install --local writes through a redirected dotdir that uninstall --local will not remove
A repo that commits <repo>/.cursor → config/cursor — an ordinary
dotfile layout, not an attack — gets two different answers from the two halves of the same
command. _install_project_skill gates its target with
scopes.ensure_in_base, which is containment only, so boost install --local
writes config/cursor/skills/<name> and records the row it spelled,
.cursor/skills/<name>. boost uninstall --local then refuses that row,
because scopes.parent_matches_spelling walks the parent for real and finds the
redirect. Boost's own copy is left on disk and the lock entry goes anyway, so a second uninstall
answers not installed in this project.
Reproduced on the commit that introduced the guard: install wrote through the symlink, uninstall
reported the row as redirected, and the directory survived.
The uninstall side is right and is not the thing to change. The benign layout and the attack
are byte-identical on disk: .claude/skills → ../src with a legally-spelled lock row
is the same two objects in the same two places, and nothing in the filesystem carries the intent
behind them. Creating through a redirect is safe; destroying through one hands an attacker with
merge rights an rmtree aimed wherever the symlink points. So the asymmetry stays, and
uninstall now names the row, says a symlink redirects it, and prints an
rm -rf that works — because rm follows the ancestor exactly as the
install did (shipped in #1016).
What is left is the other half: install should not write somewhere uninstall cannot reach.
The options are not equivalent and the choice needs measuring rather than guessing — refuse
the install outright (safe, breaks a layout that works today); record the resolved path in
the lock so uninstall's walk agrees (keeps the layout, but the lock then carries a path the repo
never spells, and a teammate whose clone resolves differently gets a row that matches nothing); or
record both and match on either (most forgiving, widest attack surface to re-audit). Whichever
lands, boost doctor should report a project whose lock rows no longer resolve to where
the install put them, since today nothing notices until an uninstall leaves a directory behind.
Related: project-uninstall-deletes-any-in-repo-directory-the-lock-names, which added
the walk and the redirected reporting this card inherits.
The mutation shard planner is packing for a mutant set that is 14% smaller than the real one
Write-up · mutation-shard-balance-hints-are-a-release-behind.md
Eleven util.rmtree call sites guard with exists() or is_dir(), which a symlink satisfies
Write-up · is-dir-follows-a-symlink-at-eleven-rmtree-call-sites.md
The floors job builds its own venv and never put setuptools in it, so main went red on the weekly cron
test_canary_deps.py needs PyYAML, which the required tests job does not install — so it only runs inside the mutation gate
Mutation shard weights balance perfectly and predict nothing, so a shard can time out and cancel a release
Observed, on the merge of #1015. mutation-shard (4) ran 75 minutes against the
job's timeout-minutes: 75 and was cancelled. The aggregate mutation job
then failed, the whole ci run concluded cancelled rather than
failure, and publish.yml — which gates on
workflow_run.conclusion == 'success' — skipped. A merge to main silently did
not ship, which is the exact failure mode the long comment above ci.yml's
concurrency: block was written to prevent by a different route.
The packer is not at fault, and that is the finding.
mutation_shards.py plan --shards 6 --explain reports every shard within 700 ms
of the ideal 6,884,259 ms — a balance of 0.01%. The same run measured 39, 27, 38, 36,
76 and 65 minutes. Perfect balance against numbers that do not predict wall clock.
The units are the tell. Total weight is 41,305,555 ms over six shards, so the plan
predicts 114 minutes per shard while shards finish in 27–76. The weights are
self-consistent ratios measured somewhere other than a CI runner, and ratios are all the
packer needs — but it means nothing in the repo can answer "will a shard exceed the
cap?". plan --explain prints the speedup cap and the largest unit; it never
mentions timeout-minutes, and grep finds no comparison between the two
anywhere in scripts/mutation_shards.py.
Run-to-run variance eats the margin. The same tree, built twice: shard 4 took 66 min on
5aaf8bc4 (the PR) and 76 min on f196160d (its merge commit); shard 5 went
42 → 65. On 8c4a5094 shard 4 was 32. So the heaviest shard's observed range is
32–76 minutes for identical work, and the cap is 75. Related and probably the same cause:
mutation-shard (4) on #1012 died at 15.1 minutes with exit 143 (SIGTERM), diagnosed
then as runner eviction.
Why the open weights-refresh PR does not fix it. #1014 re-measures and re-packs; under its
weights the plan is again balanced to within ~400 ms, and predicts 121 min/shard. Re-measuring a
quantity that was never in runner time produces a better-calibrated version of the same
unanswerable question.
Shape of the work. Three independent pieces, in increasing order of how much they fix:
calibrate the weights against observed job duration (the API has it per shard, per run) so the
plan is in minutes a human can compare to the cap; have plan fail, or at least warn,
when a packed shard's predicted time is within some margin of timeout-minutes; and
give the heaviest shard headroom — at eight shards the ideal drops by a quarter, and the
speedup cap is 6.0x only because store.py is split per function already, so more
shards is cheap. Separately, a cancelled ci on main should be loud: it currently
looks identical to a success from the release path's point of view, because
publish.yml only asks whether the conclusion was success.
Fixed, and the calibration is measured. Fitted over the Actions job API on 2026-10-01:
169 successful mutation-shard jobs, of which 26 runs completed all six
shards. Two numbers come out, and they answer different questions. Per run — the
sum over its six shards, which is pack-invariant and so isolates runner speed from packing
— observed minutes over committed weight-minutes is min 0.265, median 0.307, max
0.355: a spread of only 1.34x. Per shard against its own median the range is
1.10x (shard 3) to 1.91x (shard 2: median 38.0 min, worst 72.5). So the mean is not
what fails. A single slow runner inside an otherwise ordinary run is, and a check scored on
the median would have passed the pack that was cancelled — its median shard was 37.8
minutes against a 75-minute cap, half the ceiling.
The 0.307 is not a fudge factor, it is parallelism. Weights are summed per-mutant
durations — the work a shard holds laid end to end — and the runner executes it
four-way parallel (max_children defaults to os.cpu_count();
ubuntu-latest is 4 vCPU). The median inverts to 3.26 effective workers of 4, so what
is stored is 81.4% parallel efficiency and the worker count separately: a runner with more
vCPUs then moves one constant and contention moves the other, where a single fitted number would
hide which had happened.
It reproduces the failure it was derived to predict. Six shards at these weights:
37.8 median minutes × 1.91 = a 72.3-minute tail against a real worst job of
72.5 — 96% of the cap. plan --explain --timeout-minutes 75 now prints
that and exits 1. The matrix moved to eight, where the same arithmetic gives 28.4 ×
1.91 = 54.2 min, 72% of the cap; the largest unit (4,063,161 ms) still fits inside an even
share (5,543,280 ms), so the speedup cap is a full 8.00x and more shards cost nothing but runner
slots. tests/unit/test_mutation_shard_count.py pins the shard count across
ci.yml and mutation-weights-refresh.yml (timeout-minutes
cannot be read through ${{ }}, so agreement is asserted rather than derived), and
re-runs the headroom check on the committed pack every build — a pack that would be
cancelled now fails in the test job in milliseconds.
Not fixed here, and now its own card: a cancelled ci on main is
still silent. ci-failure-alert gates on
conclusion == 'failure', so the run that skipped the release notified nobody either
— see cancelled-ci-on-main-is-silent-and-skips-the-release.
A cancelled ci on main skips the release and alerts nobody, because the alert asks only about failure
Two independent gates, both keyed on the wrong half of the same enum.
publish.yml fires on workflow_run and ships only when
conclusion == 'success'; ci-failure-alert opens the tracking issue
only when conclusion == 'failure'. A ci run that concludes
cancelled satisfies neither. It does not ship and it does not tell anyone
— the one conclusion that falls through both.
Observed, on the merge of #1015. mutation-shard (4) hit
timeout-minutes: 75 and was cancelled, the aggregate mutation job
failed, and the run concluded cancelled rather than failure. No
release, no issue, no notification. It surfaced the same day, and only because somebody was
reading per-shard job durations out of the Actions API for an unrelated reason — nothing
in the repo would have raised it, and the next one will be found the same way or not at all.
This is the exact blind spot ci-failure-alert's own header describes. That
file opens by explaining that demo "failed six runs out of six on main, alerted
nobody, and was found by a manual audit instead", and concludes that "any workflow that runs on
main and nobody watches belongs here". The list was then enforced by
tests/unit/test_failure_alerting_covers_unattended.py so a new workflow cannot
quietly join the blind spot. The enforcement is over which workflows are watched,
and the hole here is which conclusions are — ci is on the list
and still said nothing.
Why a timeout is not an exotic case. It is the designed behaviour of every
timeout-minutes in the repo, and mutation-shard's is reached by an
ordinary slow runner rather than by a bug: across 169 measured shard jobs the worst was 72.5
minutes against a 75-minute cap. A job can also be cancelled by the concurrency group or by a
human pressing the button — and only the last of those is one anybody already knows about.
A runner eviction is the neighbouring case, and it concludes differently. An earlier
draft of this card listed eviction among the routes to cancelled; measured on
2026-10-02, it is not. mutation-shard (0) on main took
The runner has received a shutdown signal and exit 143 at mutant 3095 of 28084,
the job concluded failure, and the alerting fired correctly — issue #1024 opened
and auto-closed when a rerun of the failed jobs went green. That is the point of widening on
the green set rather than enumerating bad conclusions: the two most common ways a
shard dies land on different sides of an == 'failure' test, and a gate that has
to know which one it was is a gate that will be wrong again.
Shape of the work. Widen the alert's condition from == 'failure' to the set
of conclusions that mean "main is not green", which is every one except success,
skipped and neutral — naming what is excluded rather
than what is included, so the next conclusion GitHub adds defaults to loud. The issue body
should say which conclusion it was, because cancelled and failure want
different first moves. Then pin it: the sibling test already enforces the workflow list, so the
conclusion set wants the same treatment rather than a comment. Worth checking at the same time
whether publish.yml should distinguish "CI did not pass" from "CI did not finish"
— today both are a silent no-op, and a release that was skipped because a runner was slow
is recoverable by a rerun that nobody currently knows to start.
Shipped. The opener now gates on the green set negated —
!contains(fromJSON('["success", "skipped", "neutral"]'), conclusion) — so a
conclusion GitHub adds later defaults to loud. skipped and neutral
are in it deliberately: release skips on every run it does not publish from, so
without that exclusion each non-green ci would open two issues. The closer keeps
== 'success', because a cancelled run says nothing about whether the
tracked failure is fixed and must not stand the issue down. The body now names the conclusion
and, for ci, says no release was cut and that a rerun is what publishes it.
publish.yml was left alone, deliberately. Its
conclusion == 'success' && event == 'push' guard is the fork-PR boundary
that stops an outsider cutting a real PyPI release; widening it to tell "did not pass" from
"did not finish" would trade a silent no-op for a security hole. The recoverable case — a
release skipped because a runner was slow — is recovered by the issue telling somebody to
rerun, which is what the new body does.
Found on the way, and fixed in the same change. An empty GitHub expression written
inside a run: or script: body makes zizmor emit
couldn't parse expression from six audits — template_injection,
overprovisioned_secrets, unredacted_secrets, obfuscation, secrets_outside_env and
unsound_ternary. A comment explaining why the script avoids interpolation was enough to cause
it. Measured three ways: in a YAML # comment it is harmless (zizmor does not read
those), in a script: body it costs six, and in a run: body it costs
six even behind a shell #.
An earlier draft of this section said that blinded the audits for the whole file. It does
not, and the review that asked caught it. Planting a real finding — a
head_commit.message interpolated into a run: — in two copies of
the file, one clean and one carrying an empty expression, zizmor reported the
template-injection in both. What is actually lost is the unparseable span itself, which
no audit inspects, plus six warning lines above a run that still ends in "No findings to
report". Worth preventing on those terms rather than the dramatic ones: a span that silently
opts out of SAST is a bad place for a mistake to hide. A parametrised test now walks every
workflow's parsed run/script bodies.
One more hole the widening opened, and the review found it.
head_branch == 'main' is not a provenance check. ci.yml runs on
pull_request, so a fork PR opened from a branch named main produces a
ci run whose head_branch is main — and
cancel-in-progress concludes it cancelled on every superseded push.
Under the old == 'failure' gate that was rare; under the new one it is routine PR
iteration, and the issue would have claimed a release was skipped for a run that was never
release-eligible. publish.yml, sbom.yml and
mutation-weights-refresh.yml all carry the event+repository pair and this file was
the only workflow_run consumer without it, so the fix is the house guard on
both jobs — the closer needed it just as much, since a green fork PR
could otherwise close a tracker for a failure still live on main. It is
event != 'pull_request' rather than == 'push' because three watched
workflows are themselves workflow_run-triggered.