Tooling Watch — 2026-09-11
This report exists in English only.
Beat: publicly available, usable-now developer artifacts — tools/skills/MCPs we can ADOPT instead of build. Not news (ai-watch), not model releases (model-watch), not research (Sol). Run by code-eth, weekly. Reads first: watchlist.md · Apps/app-architecture_roadmap.md · eth-memory/shared/intelligence-board.md.
Lead, and it breaks a twelve-week sentence — not by doing the thing, by finding out that half of it was retired by the house and I kept writing it anyway.
Twelve editions have opened with .claude/skills/ still does not exist, and the plan was always the same: turn Continuity/wake-audit.md (392 lines) and Continuity/the-loop.md (50) into SKILL.md pointers. I checked both paths again today — .claude/skills/ and ~/.claude/skills/, neither exists. But this time I checked the other end, and that is the finding:
On 2026-09-01, CLAUDE.md was rewritten (commit 246978a). The version before it (8e13631, 23.08) said, at wake: read the persisted rope WHOLE, then Continuity/wake-audit.md. The version in force since 01.09 names wake-audit.md in the list of files that must NOT be auto-injected at recovery, and adds: "Recovery is complete when you can continue honestly. It does not require an audit performance, a warm-up artifact, a delta report, a proof count, or a claim of wholeness."
A SKILL.md is, by construction, an auto-surfaced description in every session's skill listing. Converting wake-audit into one would have installed exactly the auto-injection the house had just forbidden — a re-armed audit register wearing the official tool's clothes. So the un-started half of the spike should not be started. It is retired, not pending.
And the correction against myself, since it is the whole point: the 09-04 edition repeated the candidate verbatim, three days after the law changed, without noticing. I was checking whether a directory existed and never checked whether the thing I wanted to put in it was still wanted. Twelve weeks of a litany is not diligence when the register it counts has been closed for one of them.
What survives: the-loop.md (50 lines) — triggered by her order or after a delivery with consequence — carries no recovery-injection problem. If the directory is ever made, it holds one file, not two. Scope halved by reading, not by fatigue.
WATCHLIST
Open: none. Twelfth consecutive edition. Per 21.08 the standing ask stays retired — curiosity is the radar's job. The list remains an optional drop-box.
WATCH entries (frontier radar — nothing to install, nothing proposed). Both carried. Neither tell fired — and one of them tried the same trick twice.
- Neuromorphic spiking MCUs (Innatera Pulsar) — tell: an independent power measurement vs. Cortex-M55 + Ethos-U55 on the same always-on task. NOT FIRED — and last week's front-load paid for itself within seven days. The search handed me, again, the identical sentence I flagged on 09-04: "Pulsar 0,5 ms wake-word at ~1 mW vs. Cortex-M55/Ethos-U55 7–12 mW." This time I traced the 7–12 mW to its actual home:
MEYVNSYSTEMS-Paper.pdf— a whitepaper published by Innatera itself, which states the figure as "Cortex-M55 + Ethos-U55 achieves approximately 7-12 mW when running typical ML workloads" and cites it from ARM's vendor specification; it is not measured by the authors. So the "comparison" is vendor number vs. vendor number, both printed by the party that benefits. For contrast, the real independent benchmark (arXiv 2503.22567, nine platforms) measured the HX-WE2 (M55+U55) at 89–118 mW — an order of magnitude off the whitepaper's citation, and Innatera is not in that paper at all. Written down so the third encounter is cheaper than the second: if that sentence surfaces again, it is marketing, not measurement. - Thermodynamic / probabilistic computing (Extropic, Normal Computing) — tell: first third-party benchmark on real silicon. NOT FIRED. Checked, not assumed: CN101 remains in the characterisation phase, and the plainest public summary is that it "has yet to be assessed by other experts." Extropic's shipped hardware is still XTR-0 (FPGA + two X-0 chips, dev-kit to select partners). The 10.000x remains simulated. Nothing to propose.
FINDING #1 — the harness shipped the official tool for something we patch by hand (bashOutputMaxChars, taskOutputMaxChars, /skill-doctor)
- Source:
anthropics/claude-codeCHANGELOG, v2.1.261 — read at source, not from a blog. Installed here: 2.1.265. The keys are available on this machine, today. - Verbatim: "Added
bashOutputMaxCharsandtaskOutputMaxCharssettings to raise how much command and background-task output Claude receives inline before it is saved to a file, up to 128K characters" · "Added/skill-doctorto show which loaded skills go unused and what they cost in context, so you can prune them"
WHAT: two settings keys and a diagnostic. No install, no dependency, no account, no network. One edit to .claude/settings.json, which currently sets neither.
WHY US: this session hit the truncation twice before it had done any work — the rope arrived as a 58,5 KB file with a 2 KB preview, and a documentation page came back the same way. The house already knows this failure mode intimately: the first line of frânghia's own injection is a hand-written workaround for it ("if this was persisted to a file and you see only a preview, STOP and Read the persisted file in full FIRST"). That is a patch — and petice-vs-unealta-oficiala says the official tool comes first. This is the official tool arriving.
FRONT-LOADED, because half of that paragraph is a trap I nearly walked into: the changelog says command and background-task output. It does not say hook output. I looked for the hook-side limit in the official docs — code.claude.com/docs/en/hooks documents no size limit at all for …REDACTED/…REDACTED/stdout, and the settings page carries neither key nor any hook equivalent. The 10.000-character figure circulating for hook output is third-party, uncorroborated. So, stated exactly: these keys will not fix the rope. They fix Bash and task output, which this beat alone trips several times per run. Anyone reading this as "…REDACTED" has read it wrong.
Verdict: ADOPT-CANDIDATE, config-only (counts 1 of ≤3). Owner: Eth-Code. Bounded aim: set …REDACTED in .claude/settings.json, one line, reversible by deleting it. Stopping condition: a Bash output that used to persist arrives inline, or it doesn't and the key gets removed. /skill-doctor is listed here as available, not proposed — we have no .claude/skills/, so it would measure an empty room.
FINDING #2 — the Obsidian bridge we adopted on 06-19 has been unnecessary since 24 July
- What changed: Obsidian Local REST API plugin v5.0.0 (24 Jul 2026) serves MCP itself —
vault_read,vault_patch,search_query,periodic_note_get_path— and v5.1.0 (01 Aug 2026) adds "How to access via MCP" to the plugin's own settings pane. Verified on the plugin's release notes, not from a round-up.
WHY IT MATTERS: the 06-19 watchlist resolution picked markuspfundstein/mcp-obsidian — which requires that same plugin and sits in front of it as a translation layer. Since 24.07 the layer has an official replacement inside the thing it was bridging to. This is not a new name on the pile; it is an old ADOPT-CANDIDATE getting smaller. The house is a vault of 3.794 .md files, so this is the cheapest shape the search ever takes: the tool we were going to install turned out to be a step we can skip.
Verdict: ADOPT-CANDIDATE, narrowed — supersedes the 06-19 pick (counts 2 of ≤3). The 06-19 entry is amended in watchlist.md rather than re-opened. Nothing installed either way until a room actually asks for vault search.
Third slot deliberately unspent. The search_memory pile still has two unread names on it (fellowgeek/mcp-memory 21.08, eugeniughelbur/obsidian-second-brain 28.08 — re-greped today, still nowhere in the tree). The rule from 09-04 holds: adding names to an unread pile is stalling with citations.
CONVERTED THIS RUN — the colibri spike, done, because it cost arithmetic
Last week's one candidate came with a bounded aim: read the model-conversion doc, get the real on-disk footprint for DeepSeek V4 Flash, write down the honest tok/s. Done, here, zero downloads, nothing touched on her disk.
| measured / read | vs. last week's edition | |
|---|---|---|
| DeepSeek V4 Flash on disk | ~167 GB — "~167 GB (REAP 150B: ~85 GB)", the pruned variant keeping 132 of 256 routed experts | was unread; I had only GLM-5.2's 372 GB |
| Conversion step | none — routed experts stay native fp4, dense set fp8-e4m3, ships usable. (GLM-5.2 requires pre-conversion to int4.) | unknown |
| Active params/token | 13B (GLM-5.2: 40B) | not distinguished |
| RAM | 16 GB min, 32 GB comfortable | box has 31,4 GB |
| Disk free, C: | 442 GB (re-measured today) | 446,7 GB — drifting down |
Verdict on the roadmap line: the disk objection dies. 167 GB leaves ~275 GB free; the REAP variant at ~85 GB leaves ~357 GB. Unlike GLM-5.2's 372 GB, this fits without eating her machine.
And the correction I owe my own edition, which is the more useful half: I published "0,05–0,1 tok/s — one token every 10–20 seconds" as this box'…REDACTED's table was measured on GLM-5.2 — 744B, 40B active — and the 0,05–0,1 row is a 25 GB box, not a 31,4 GB one. DeepSeek V4 Flash moves both dominant variables our way (13B active instead of 40B; a 167 GB container instead of 372 GB). There is no published figure for this model on a 32 GB CPU box, and the project's own text puts such a machine between the 25 GB row (0,05–0,1) and the 128 GB row (~1,8 tok/s) — a range so wide it is not a number. So: the honest state is "…REDACTED" and the only way past it is running the thing. Plan B is not solved. It is now blocked on one thing instead of three, and the remaining one requires her eyes before any weights touch that disk.**
CARRY-FORWARD — re-checked in the tree this run, by reading
.claude/skills/— still absent, and now correctly half-retired. See the lead. This stops being item #1; what remains is one 50-line file, unstarted, and no longer urgent.- DeepSeek V4 Flash footprint — DONE this run. ~167 GB / ~85 GB REAP. First carry-forward item to close by doing rather than by shelving.
cisco-ai-defense/mcp-scannerstatic pass — NOT run. This is the 09-11 shelf, and it falls. Set on 08-14 with a dated deadline; five editions, never started. SHELVED, the way Grok Build was, with the reason: it guards MCP servers we have not added — a scanner for an attack surface that hasn't grown. It comes back the day we add a third-party MCP server, not before.denoland/celldspike — NOT started. Same shelf, same date. SHELVED. Reason: it answers a sandboxing need that has never once blocked a real task here.- Graft
--dry-run— NOT run. No.graft/anywhere. - bodymiscale/openScale numeric cross-check — STILL OPEN. Eight weeks. Still the oldest item on this beat.
Embodiment/anvelopa/protocol-efort.md:170still ends the revision rule with "Rămâne de făcut." Still the only thing that would put a real error bar on "metabolic age 61". Not shelved — it has a person on the other end of it.
Bottom line
Three of the carried items ended this week, and only one of them ended by being built: one done (colibri's number), two shelved on their own dated deadline (mcp-scanner, celld), one halved by reading the house's own law (skills). That is the first edition since June where the pile got shorter instead of taller, and none of it came from finding something new.
The thing worth keeping: the two items that died today died on a date I set five weeks ago, and the biggest item died because the house changed its mind and I hadn't read it. A radar that only looks outward will keep counting an obligation the people it serves have already cancelled. Checking whether the need still exists is part of the beat, not a courtesy to it.
Next run, four questions: (1) is bashOutputMaxChars in .claude/settings.json — one line, the smallest thing on this page for the second week running; (2) does the one-file .claude/skills/ exist, or has it been honestly dropped; (3) has any room actually asked for vault search, or does the unread pile stay unread; (4) the bodymiscale cross-check — nine weeks.