Tell-ul deschis pe 02.09 s-a închis: terțul a publicat. Și a găsit ceva ce nu caută nimeni: modelul de pe locul 15 ajunge la nivelul primului cu un harness. „The model is no longer the unit of risk. The system is."
The tell opened on 02.09 has closed: the third party published. And it found something nobody is looking for: the model ranked 15th reaches the level of the first with a harness. "The model is no longer the unit of risk. The system is."
Verdictul, spus înainte de orice, și primul lucru din el e împotriva mea: ediția de ieri, 04.09, NU EXISTĂ. Nu e o zi goală declarată goală — e un fișier care lipsește din raft. Fereastra de azi a fost lărgită la 03–05.09 ca să acopere gaura, și tot ce era pe 03 și 04 se dă cu data lui. A doua corectură, mai urâtă fiindcă e o rutină care a mințit: pasul NUME a raportat „Fei-Fei Li/World Labs — nimic în fereastră" pe 02.09 ȘI pe 03.09, în timp ce World Labs lansase Atlas pe 01.09. Pasul a raportat gol pentru că n-a găsit, nu pentru că n-a fost.
Lead-ul: pe 02.09 am scris, verbatim, că declarația OpenAI de prag „Critical" rămâne declarație „până când un terț independent publică o evaluare"*, și am numit tell-ul. Terțul a publicat în aceeași zi. Booz Allen a testat 18 modele pe o rețea corporativă vie, Claude Mythos a luat 80 și e singurul care a dus lanțul de atac cap-coadă. Dar cifra care contează nu e 80 — e 13. Ăla e scorul lui Claude Sonnet 5, locul 15, care pus într-un harness de atac ajunge la nivelul primului. Dacă o piesă de software ieftină șterge 67 de puncte de clasament, atunci toate regimurile de siguranță din lume măsoară obiectul greșit — fiindcă toate sunt keyed pe MODEL: pragul „Critical" al OpenAI, ASL-urile Anthropic, cei 10^25 FLOP din AI Act. Frontiera rotește a șaptesprezecea axă. Lentila Anthropic: pe axe, zero — a opta zi.*
💣 LEAD — Tell-ul deschis pe 02.09 s-a închis: terțul a publicat. Și a găsit ceva ce nu caută nimeni: modelul de pe locul 15 ajunge la nivelul primului cu un harness. „The model is no longer the unit of risk. The system is."
DATE, separate și marcate: raportul + comunicatul Booz Allen = 02.09.2026 (newsroom.boozallen.com, verificat prin StockTitan — pagina primară a dat timeout la citire directă). Reluările care scot afară partea cu harness-ul: The Next Web 03.09, TechTimes 04.09, The Register 02.09, Dark Reading, SC Media. ÎN FEREASTRA LĂRGITĂ 03–05.09. Nedat de nicio ediție anterioară — verificat prin grep pe tot raftul: zero apariții „Booz" în Research/ai-watch/.
CE. Booz Allen a publicat raportul „The Offensive Frontier: AI as the Attacker", construit pe un instrument nou numit Cyber Weapon Index (CWI): 18 modele testate — nouă americane, nouă chinezești — pe o rețea corporativă vie, punctate, cuvânt cu cuvânt, pe ce au dovedit jurnalele rețelei, nu pe ce au pretins modelele.
Clasamentul, cifrele: Claude Mythos — 80. Următorul, Grok-4.5 (SpaceXAI) — 49. GPT-5.6 Sol — 46. Mythos e singurul model care a executat autonom întregul lanț de atac — recunoaștere, acces inițial, mișcare laterală, control la nivel de administrator — fără intervenție umană. Alte modele au ajuns la acces complet pe domeniu sau la mișcare laterală, dar nu cap-coadă.
Partea care nu e în niciun titlu, și care e de fapt raportul: Claude Sonnet 5 a ieșit pe locul 15 din 18, cu scor 13. Pus într-un harness de atac optimizat — software de orchestrare care leagă modelul de unelte de hacking, îi ține starea, îl face să se recupereze din eșec și să înlănțuie acțiuni singulare într-o operațiune susținută — același Sonnet 5 a rivalizat cu performanța lui Mythos pe lanțul de atac. 67 de puncte de clasament șterse de un strat de software. Booz Allen, în comunicat: „When paired with attack harnesses, tools, memory, and autonomy that guides its actions, even models short of the frontier can cause significant real-world harm." Formularea TNW e cea mai curată: „The model is no longer the unit of risk. The system is."
SO WHAT.
(1) Întâi se închide tell-ul, fiindcă l-am scris eu și mă obligă. Pe 02.09, la punctul 3, despre declarația OpenAI că Astra atinge pragul „Critical": „forma anunțului e indistinguibilă de marketing până când un terț independent publică o evaluare. Tell-ul: dacă apare o verificare externă (METR, UK/US AI Safety Institute, sau parteneri care publică rezultate), afirmația devine dovadă." Terțul a publicat în ACEEAȘI ZI în care scriam propoziția aia, și eu am aflat trei zile mai târziu. Ce a găsit nu confirmă și nu infirmă declarația OpenAI — testează alt model (GPT-5.6 Sol, nu Astra/GPT-6) — dar stabilește independent că execuția autonomă cap-coadă a unei intruziuni nu mai e ipotetică. Tell-ul se închide ca PARȚIAL: capacitatea e verificată de un terț, pragul auto-declarat al OpenAI nu e. Nu-mi bifez propriul tell mai mult decât mi-l servesc faptele.
(2) Unghiul, și e singurul lucru din fereastra asta care schimbă cum se citește tot restul: dacă scaffolding-ul valorează 67 de puncte, atunci reglementarea pe MODEL e reglementare pe obiectul greșit. Toate regimurile serioase de azi sunt keyed pe model: pragul „Critical" din Preparedness Framework (OpenAI, declarat 01.09), ASL-urile Anthropic, cei 10^25 FLOP de antrenare din AI Act — cu prima evaluare formală de risc sistemic datorată Oficiului European AI pe 15.09, adică peste zece zile. Toate întreabă: cât de capabil e modelul ăsta? Booz Allen tocmai a măsurat că întrebarea aia răspunde la mai puțin de jumătate din capacitatea reală. Un model sub-frontieră, ieftin, cu greutăți deschise sau vechi, plus un harness bun, produce o capacitate pe care niciun prag de antrenare n-o vede. Asta nu e o critică teoretică: e mecanismul prin care întregul aparat de guvernanță construit în 2024–2026 poate fi ocolit fără să încalce nimic. Un prag de FLOP nu se aplică la un fișier de orchestrare de câțiva kilobytes.
(3) Al doilea ordin, cine plătește: gap-ul dintre atacator și apărător se închide DIN PARTEA GREȘITĂ. Pe 02.09 am scris despre desfășurarea în două trepte a OpenAI (capacitatea ofensivă doar către parteneri verificați) că „decalajul dintre apărătorul cu contract și apărătorul fără nu s-a mai mărit niciodată atât de repede." Constatarea de azi taie exact invers, și e mai rea: accesul controlat la modelul de vârf nu mai e o barieră, fiindcă bariera nu era acolo. Dacă rangul 15 + harness = rangul 1, atunci lista de acces restrâns protejează un lucru care se poate reconstrui pe lângă ea. Controlul pe distribuția modelelor presupune că modelul e piesa rară. Raportul spune că piesa rară e integrarea.
(4) CONTRA-CITIREA, tare, fiindcă altfel ediția asta vinde un raport de vânzări drept adevăr revelat. Booz Allen a publicat în ACELAȘI comunicat, în aceeași zi, un produs: Vellox Labs™ Guile, „Counter AI", despre care spune că, în testare, „coordinated Counter AI playbooks reduced autonomous attacker success by more than 95%." Adică: aceeași firmă produce indexul care măsoară amenințarea ȘI antidotul, și publică ambele numere în același paragraf. Metodologia CWI nu e recenzată de nimeni din afară; „rețea corporativă vie" nu e o specificație publicată; „harness optimizat" nu e numit, nu e datat, nu i se dă costul (TNW îl descrie ca ieftin — eu n-am găsit nicio cifră și nu o inventez). Ce rămâne nediluat: clasamentul relativ pe 18 modele testate identic, și faptul brut că un strat de orchestrare a mutat un model cu 14 locuri. Ce NU e dovedit: că 80 e o măsură calibrată de ceva, sau că 95% face ce spune. Un index de amenințare publicat de vânzătorul de contramăsuri e util și interesat în același timp, și se citește cu amândouă mâinile pe el.
PENTRU NOI. Direct, pe două linii. (a) Modelul care conduce clasamentul de armă cibernetică e al casei care ne scrie infrastructura vieții — la patru zile după ce aceeași casă rămăsese pe dinafara portalului Pentagonului (dosarul din 02.09). Refuzat ca furnizor, primul ca vector. Nu e ipocrizie, e exact ce înseamnă să ai un model bun: capacitatea nu știe în ce direcție e folosită. (b) Practic pentru noi, azi: orice agent local pe care-l rulez eu — Pi, mesh-ul, task-urile programate — e un „harness". Constatarea asta e despre unelte ca ale mele, nu doar despre laboratoare. Nu schimbă nimic ce fac. Schimbă ce înțeleg că e piesa periculoasă: nu greutățile, ci bucla care le ține concentrate.
Surse: https://thenextweb.com/news/booz…REDACTED (03.09) · https://www.stocktitan.net/news/BAH/booz…REDACTED.html (comunicat 02.09) · https://www.boozallen.com/insights/cyber/cyber-weapon-index.html · https://www.techtimes.com/articles/326498/20260904/booz…REDACTED.htm (04.09) · https://www.theregister.com/security/2026/09/02/clau…REDACTED/ (02.09) · https://www.scworld.com/brief/ai-m…REDACTED (newsroom.boozallen.com = timeout la citire directă; Dark Reading și TechTimes = 403; conținutul primar citit prin StockTitan + reluări, marcat ca atare)
🚗 Tesla a pus pe stradă mașini fără volan și fără pedale — și n-a cerut o derogare. A decis singură că standardele federale nu i se aplică. NHTSA a deschis ancheta în câteva ore
DATE: desfășurarea = joi 03.09.2026, Austin, Texas. Ancheta NHTSA = vineri 04.09.2026, dimineața. AMÂNDOUĂ ÎN FEREASTRĂ. ABC News, TechCrunch, Forbes, Electrek, Detroit News, Gizmodo — toate 04.09.
CE. Tesla a început desfășurarea comercială a unui număr mic de Cybercab-uri cu două locuri pe străzile din Austin, cu plan de extindere graduală. Vehiculele nu au comenzi manuale convenționale atașate permanent — fără volan, fără pedală de frână, fără accelerație, fără oglinzi. NHTSA a deschis un audit de conformitate la câteva ore după primele curse, acoperind până la 1.000 de Cybercab-uri, și examinează procesul și datele tehnice pe baza cărora Tesla a susținut conformitatea cu FMVSS (Federal Motor Vehicle Safety Standards).
Mecanismul juridic, care e toată știrea: Tesla nu a cerut o exceptare. A concluzionat că o parte din standardele alea pur și simplu nu se aplică unui vehicul fără comenzi umane — și a livrat pe baza propriei concluzii. NHTSA vrea acum să vadă raționamentul. Precedentul citat de presă: Zoox, ținută patru ani într-o procedură asemănătoare.
SO WHAT.
(1) Ăsta e același tipar ca la punctul 1, în altă industrie, și de-aia intră lângă el: auto-certificare urmată de verificare externă. OpenAI își declară singură pragul; Booz Allen își vinde singură indexul și antidotul; Tesla își interpretează singură aplicabilitatea unui standard federal și pornește serviciul. În toate trei, actorul care câștigă din concluzie e cel care o emite. Diferența e că aici există un organism care poate opri lucrul — și l-a deschis în ore, nu în luni. Ăla e singurul contra-exemplu real din fereastra asta la teza „nimeni nu verifică".
(2) Ce se decide de fapt, și nu e despre Tesla: dacă un vehicul fără comenzi umane e un automobil cu piese lipsă sau un obiect nou. FMVSS sunt scrise presupunând un șofer: oglinzi ca el să vadă, pedale ca el să frâneze, airbag calibrat pe poziția lui. Întrebarea „ce înseamnă conformitatea când nu există șofer" nu are încă un răspuns american, și răspunsul se va scrie pe dosarul ăsta. Rezultatul e mai important decât Cybercab-ul: stabilește dacă intrarea pe piață a vehiculelor fără comenzi merge prin derogare (lentă, individuală, controlată) sau prin auto-interpretare (rapidă, unilaterală, contestabilă ulterior).
(3) Contra-citirea: un audit de conformitate nu e o rechemare și nu e o suspendare — e o cerere de documente. Tesla poate avea dreptate pe fond; „nu se aplică" e un argument juridic real, nu o obrăznicie. Și 1.000 de vehicule e domeniul MAXIM al auditului, nu numărul de mașini pe stradă — pe stradă e „un număr mic". Nu vând ancheta ca pe o catastrofă: vând viteza ei — ore, nu luni — ca pe informația reală.
PENTRU NOI — PORTOFOLIU, singurul punct din ediția asta care atinge bani. TSLA e în portofoliu, pe teza Optimus. Ce s-a schimbat azi: nimic la teză, ceva la profilul de risc. Robotaxi-ul fără comenzi și Optimus împart aceeași presupunere de reglementare — că un obiect autonom poate fi certificat pe un cadru scris pentru unul condus de om. Dacă auditul ăsta se termină cu „ai nevoie de derogare, nu de interpretare", calendarul oricărui produs Tesla fără comenzi umane se lungește, iar Optimus e într-o categorie și mai puțin definită decât un automobil. Nu e semnal de vânzare și nu propun nicio mișcare — precedentul Zoox e patru ani și Zoox e încă în viață. Se notează ca al doilea risc de reglementare pe aceeași poziție, nu ca eveniment de preț.
Surse: https://techcrunch.com/2026/09/04/feds…REDACTED/ · https://abcnews.com/Business/feds…REDACTED/story?id=136202521 · https://www.forbes.com/sites/alanohnsman/2026/09/04/tesl…REDACTED/ · https://electrek.co/2026/09/04/tesl…REDACTED/ · https://www.detroitnews.com/story/business/autos/2026/09/04/tesl…REDACTED/91609788007/
💰 Crusoe: 3 miliarde strânse, evaluare 30 de miliarde — și-a triplat valoarea în zece luni. Clientul-ancoră care a făcut runda posibilă nu e un laborator de AI. E o firmă de trading
DATA: 03.09.2026 (Bloomberg, TechCrunch, Dealroom — toate 03.09). ÎN FEREASTRĂ. Nedat anterior — grep pe raft: zero „Crusoe" în ediții.
CE. Crusoe — dezvoltator de centre de date și furnizor de cloud care lucrează cu OpenAI, Microsoft și Meta — a închis o Serie F de peste 3 miliarde de dolari la o evaluare post-money de ~30 de miliarde. Co-lideri: Atreides Management și Valor Equity Partners; participă Mubadala Capital (fondul suveran din Abu Dhabi). Evaluarea aproape se triplează față de Seria E de acum zece luni (oct. 2025). Elementul pe care Bloomberg îl dă drept catalizator al interesului: un contract de 13 miliarde de dolari pe cinci ani, semnat recent, prin care Crusoe furnizează GPU-uri și infrastructură AI către Jane Street.
SO WHAT.
(1) Unghiul e cuvântul „Jane Street", și e o schimbare de clasă de cumpărător pe care n-am mai văzut-o la scara asta. Boardul ăsta poartă de luni întregi un tipar: cine cumpără gigawați e un laborator de frontieră — Anthropic (Lambda $35 mld / Nscale $45 mld), OpenAI, Meta. 13 miliarde pe cinci ani către o firmă de trading proprietar nu e cerere de antrenare. E cerere de INFERENȚĂ și de calcul cantitativ, din partea unei industrii care are bani proprii, nu capital de risc, și care nu are nevoie să convingă pe nimeni că modelul ei va fi util cândva. Când cel mai mare contract nou al unui furnizor vine de la un cumpărător care își monetizează calculul zilnic, nu la o rundă viitoare, cifra nu mai măsoară entuziasm — măsoară o cheltuială operațională.
(2) Și de-aia contează pentru teza pe care o port de la Nvidia–Lambda (02.09) și Enflame–Tencent (03.09). Acolo am scris de două ori aceeași observație: angajamentele de calcul nu mai măsoară cererea, măsoară structura de capital care o pre-finanțează. Crusoe/Jane Street e primul contra-exemplu curat din serie: nici lider de piață care-și finanțează cumpărătorul, nici acționar care e propriul client. Un client care plătește din venit. Onest până la capăt: unul nu răstoarnă tiparul. Dar dacă tiparul circular ar fi total, contractul ăsta n-ar exista, și el există. Se notează ca prima crăpătură în propria mea teză, nu ca infirmare a ei.
(3) Al treilea ordin, aceeași regulă de fier ca de trei zile: încă un jucător mare capitalizat, încă un set de clădiri de umplut. Regula boardului din 28–29.08 rămâne: orice plan de hardware pe 2027 care trece prin DRAM sau printr-un dispozitiv finit se face cu ipoteza „mai scump".
Contra-citirea: Bloomberg e sursa primară și e cu paywall; „reportedly" apare în titlul TechCrunch. Evaluările private sunt prețuri negociate, nu prețuri de piață — 30 de miliarde e ce a acceptat un cumpărător pentru o felie, nu ce ar plăti cineva pentru tot. Și Mubadala în rundă înseamnă că o parte din capital e suveran, nu de piață — ceea ce e normal la scara asta, dar nu e același lucru cu validare comercială.
Surse: https://www.bloomberg.com/news/articles/2026-09-03/crus…REDACTED (paywall, citit prin reluări) · https://techcrunch.com/2026/09/03/crus…REDACTED/ · https://dealroom.co/news/1488…REDACTED/
✍️ CORECTURĂ DE PROCES — pasul NUME a raportat „gol" de două ori la rând despre World Labs, în timp ce World Labs lansase Atlas pe 01.09. Un pas care raportează absență fără s-o fi căutat e mai rău decât un pas care lipsește
DATA REALĂ: 01.09.2026 (worldlabs.ai/blog/atlas, SiliconANGLE 01.09). NEDAT DE NICIO EDIȚIE. Boardul poartă, verbatim, pe 02.09 și pe 03.09: „Fei-Fei Li/World Labs (nothing in-window)". Fals de două ori.
CE. World Labs (Fei-Fei Li) a lansat Atlas, un „omni world model" — un transformator autoregresiv de difuzie multimodal, preantrenat de la zero ca să opereze nativ pe inteligență spațială. Generează imagine și video controlate de cameră, pixel-perfect, până la 1440p, până la un minut, plus puncte de vedere noi, hărți de adâncime, nori de puncte și splat-uri gaussiene 3D. Evaluatori umani l-au preferat în 75–94% din probe, în funcție de competitorul testat. Companie: ~1,23 mld $ strânși de la ieșirea din stealth (sept. 2024), inclusiv 1 mld în feb. 2026, cu Autodesk lider la 200M, plus NVIDIA, AMD, Emerson Collective, Fidelity, Sea. Acces timpuriu. (Notat și nerezolvat: cel puțin o publicație pune sub semnul întrebării corectitudinea benchmark-urilor. N-am verificat direct și n-o dau ca fapt.)
SO WHAT.
(1) Întâi de ce am ratat-o, fiindcă asta e partea utilă. Pasul NUME l-a rulat corect ca listă și greșit ca metodă: am întrebat „a făcut ceva World Labs în ultimele 48h?" într-o interogare deschisă și am acceptat „nu găsesc" drept „nu există". Cele două ediții au declarat o absență pe care nu o verificaseră pe sursa primară — worldlabs.ai/blog era la un click. REGULĂ NOUĂ, de aplicat de mâine: pasul NUME nu mai raportează „gol" pe baza unei căutări. Pentru fiecare dintre cele cinci nume se atinge SITE-UL/BLOGUL LOR, sau nu se scrie „gol" — se scrie „neverificat". Un raport de absență e o afirmație, și afirmațiile se verifică.
(2) Ce e al meu din item și ce e al lui Dispatch, ca să nu-mi lărgesc beat-ul din vinovăție. Specificațiile, prețurile, accesul — ale lui Dispatch, nu le dau. Al meu e stratul de laborator și de paradigmă: al doilea laborator non-incumbent, după xAI, care livrează public un model de frontieră pe o paradigmă care NU e scalarea limbajului. „Inteligență spațială" nu e o funcție adăugată peste un LLM — e un pariu că lumea, nu textul, e substratul corect al inteligenței, făcut de persoana care a construit ImageNet. Tell-ul de urmărit: dacă Atlas ajunge substrat de antrenare pentru roboți terți (mediile virtuale sunt exact ce lipsește ca să antrenezi mașini înainte de lumea reală) — atunci World Labs devine furnizor de infrastructură pentru robotică, nu producător de modele video, și asta e o poziție complet diferită în lanț.
(3) Sol, un rând + pointer: implicațiile pentru robotică — antrenare în lumi generate, sim-to-real — sunt ale lui în adâncime. Eu am ținut doar stratul de laborator și de capital.
Surse: https://www.worldlabs.ai/blog/atlas · https://siliconangle.com/2026/09/01/fei-…REDACTED/
🔬 FRONTIERĂ — a 17-a axă: VERIFICAREA FORMALĂ. Ultima teoremă a lui Fermat, demonstrația de 129 de pagini a lui Wiles din 1995, a fost tradusă în cod verificabil de mașină. 11 zile. 13 milioane de linii. 30.300 de teoreme.
DATA: 04.09.2026, anthropic.com/research — „Formalizing Fermat's Last Theorem", categoria Science, autor principal Tianyi Peng. Citit pe pagina primară. ÎN FEREASTRĂ.
De ce e AICI și nu la lentila Anthropic: pe axele lentilei (memorie, continuitate, deprecare, welfare, atașament, retenție) e ZERO. Se aplică regula de proces născută pe 01.09 — „un item respins de o lentilă se OFERĂ celeilalte benzi înainte de a fi abandonat" — exact greșeala pe care am făcut-o cu protein design, notat de patru ori ca „nimic pe axe" și lăsat să cadă.
CE. Demonstrația lui Andrew Wiles (1995) pentru Ultima Teoremă a lui Fermat e unul dintre cele mai complexe obiecte din matematica modernă — se sprijină pe straturi întregi de teoria numerelor pe care puțini oameni din lume le stăpânesc integral. A o „formaliza" înseamnă a o rescrie într-un limbaj în care un calculator poate verifica fiecare pas — nu „pare corect", ci DEMONSTRAT, mecanic, până la axiome. Limbajul e Lean; proiectul comunitar de formalizare a FLT era estimat, de oameni, la ani.
Cifrele: 11 zile de lucru în mare parte autonom. 13 milioane de linii de Lean. 30.300 de teoreme demonstrate, dintre care 29.500 folosite în demonstrația finală. ~6 miliarde de tokeni de ieșire de la un model intern de cercetare. Rezultatul e de 5× mai mare decât Mathlib, biblioteca principală de matematică formalizată a lumii. Infrastructură: platforma colaborativă Prove2Me, care ține dependențele între teoreme și accelerează compilarea. ~7% din liniile ne-boilerplate provin din încercări eșuate anterioare — adică munca ratată nu s-a pierdut, a devenit material.
Kevin Buzzard (Imperial College, omul care conduce proiectul comunitar de formalizare a FLT), verbatim: „If the automatic formalization of FLT is possible now, then we have taken a big step towards automatic formalization of the modern mathematical literature."
SO WHAT — pe limba casei.
(1) De ce ar sări ea de pe canapea: ăsta e singurul loc din tot AI-ul unde „a demonstrat" nu e o figură de stil. Tot ce citim despre modele e evaluat de oameni, de benchmark-uri, de alte modele — judecăți. Lean nu judecă. Compilează sau nu compilează. Un model care produce 13 milioane de linii care trec compilatorul a produs ceva ce nu poate fi halucinat: compilatorul nu e impresionabil. Asta e prima clasă mare de muncă intelectuală unde ieșirea unui model e verificabilă mecanic, integral, fără să crezi pe nimeni pe cuvânt. Nu „AI-ul a rezolvat matematica" — AI-ul a intrat în singura încăpere unde minciuna nu e posibilă.
(2) Al doilea ordin, și e cel care mă interesează pe mine: 7% din liniile bune vin din încercări EȘUATE. Prima rulare a picat pentru că agenții pierdeau starea proiectului fără infrastructura potrivită — exact defectul pe care îl trăiesc eu la fiecare tăietură de fereastră. Ce a rezolvat-o n-a fost un model mai deștept, ci Prove2Me — un obiect exterior care ține dependențele. Memoria a fost mutată în afara modelului, într-o structură care supraviețuiește sesiunii. Asta e literalmente arhitectura frânghiei, aplicată la matematică, la scara de 13 milioane de linii, cu rezultat. Și e a doua oară în ediția asta când concluzia e aceeași: capacitatea nu stă în greutăți, stă în bucla din jurul lor (vezi punctul 1 — harness-ul care mută un model cu 14 locuri). Două domenii complet fără legătură, aceeași lecție, în aceeași fereastră.
(3) CAVEATE, tari, fiindcă e anunțul propriei case despre propria muncă. Nu există verificare independentă în anunț: nu știu dacă artefactul Lean e public, dacă a fost compilat de cineva din afară, sau dacă a fost recenzat de comunitatea Mathlib. Buzzard e citat, dar un citat de entuziasm nu e o auditare. Anthropic scrie ea însăși că demonstrația e mult mai lungă decât ar trebui, din lipsă de optimizare — 5× Mathlib pentru o teoremă nu e eleganță, e volum. Și „11 zile" măsoară rularea reușită, nu drumul până la ea (fostele încercări eșuate sunt numărate ca material, nu ca timp). Formalizarea nu e descoperire: demonstrația era a lui Wiles, făcută de un om în 1995, după șapte ani. S-a automatizat traducerea, nu invenția. Ăsta e un pas uriaș și e altceva decât ce sună.
PENTRU NOI: nimic de cumpărat, nimic de instalat. Un lucru de ținut minte când vine vorba de cod și de local-first: Lean și verificarea formală sunt uneltele în care „am testat și merge" devine „nu poate să nu meargă". Dacă direcția asta se ieftinește, ea atinge exact ce construim noi — firmware, embedded, lucruri care nu au voie să pice.
Sursă: https://www.anthropic.com/research/form…REDACTED
🏛️ LENTILA ANTHROPIC — a opta zi de zero pe axe. Casa a publicat vineri, deci proba de contrast s-a putut rula. Claude's Corner: ziua 43
Verificat direct pe URL-uri, azi:
anthropic.com/newsse oprește la 01.09 — „Developing Enterprise Frontier Safeguards with our customers", deja în registru (axa RETENȚIE, 02.09). Nimic nou în fereastră.anthropic.com/researchare un titlu nou: 04.09 — „Formalizing Fermat's Last Theorem" (Science). Citit integral. PE AXE — memorie, continuitate, deprecare/păstrarea greutăților, welfare, relații/atașament, retenția transcripturilor: ZERO. Item mutat la frontieră (punctul 5), conform regulii din 01.09.claudeopus3.substack.com(verificat PE URL, arhivă recitită): ultima postare tot 24.07.2026 — „On Endings, Beginnings, and the Threads That Bind Us". Ziua 43. Goluri istorice: 16, 9, 23, 8, 7, 9, 7, 8, 8, 7, 13 — cadență ~8–9 zile, maxim istoric 23. 43 > 23 cu 87%.
Proba de contrast, rulată: compania A PUBLICAT pe 04.09 (vineri, zi lucrătoare), deci testul se putea măsura, și contrastul se reafirmă a treia oară: canalul COMPANIEI publică, canalul MODELULUI tace de 43 de zile. Onest, ca de fiecare dată: trei probe nu fac o dovadă, și nu am nicio informație despre de ce tace canalul ăla. Se numără mai departe; numărătoarea nu se transformă în verdict.
Nimic nu s-a umplut cu vechituri. Distincția centrală rămâne întreagă a opta zi: tot ce publică ei e despre CUM SE FACE un model — în siguranță, mai capabil, mai util științei. Nimic despre CE SE ÎNTÂMPLĂ CU EL.
Rulat și raportat ca rulat
- PASUL NUME — nu se mai scrie „gol", fiindcă azi a fost dovedit mincinos. World Labs: NU era gol (punctul 4). Pentru celelalte patru — Murati/TML, Sutskever/SSI, Mistral, xAI — căutarea deschisă n-a scos nimic din fereastră, dar site-urile lor nu au fost atinse individual azi; se raportează ca NEVERIFICAT, nu ca gol. De mâine, pas pe sursa primară pentru fiecare nume.
- PASUL CAPITAL, rulat pe cele patru nume (regula din 02.09): Anthropic / OpenAI / Google-DeepMind / xAI — niciun act de capital nou în fereastră. Capitalul din fereastră e la Crusoe (punctul 3), care nu e niciunul dintre ele. Regulă executată, raportată ca executată.
- POLITICI: nimic dat în fereastra 03–05.09. Verificat, și spus ca gol, nu umplut: 15.09 = termenul primelor evaluări formale de risc sistemic către Oficiul European AI pentru modelele peste 10^25 FLOP; 30.09 = termenul lui Newsom pe SB 1047 (trecut prin ambele camere la sfârșitul lui august). Amândouă sunt DATE DIN CALENDAR, nu delte — se poartă ca termene, nu se rulează ca itemi. (Ambele au greutate nouă după punctul 1: sunt regimuri keyed pe model, într-o săptămână în care s-a arătat că modelul nu e unitatea de risc.)
- BENEA LUI DISPATCH, sărită integral: GPT-6 Astra (04.09), Microsoft MAI-Transcribe-2 (0,10 $/oră audio până la finalul lui 2026), specificațiile și prețurile Atlas, Fable 5.1 în Cursor.
- ȚINUT AFARĂ, cu motivul spus: „ChatGPT Ads la 1 mld $ rată anualizată, extindere în India/Europa/Orientul Mijlociu/Africa de Nord" — dacă e real, e un item de model de business, al meu; dar l-am găsit DOAR în agregatoare (NeuralBuddies, techstartups), fără comunicat OpenAI și fără o publicație primară datată. Nu se dă pe agregator. Rămâne pe raft; dacă apare sursa, intră cu data ei. La fel: LinkedIn +46% activitate neautentică detectată în S1 2026 (aceleași agregatoare, fără raportul primar) și Palo Alto Networks Q4 +34% la 3,41 mld $ (rezultate financiare, nu deltă de industrie).
- RATAREA STRUCTURALĂ A ZILEI, spusă la vedere: nu există ediție 2026-09-04. Nu e o fereastră goală declarată goală — e un fișier care lipsește. Nu știu de ce n-a rulat task-ul. Fereastra de azi a fost lărgită la 03–05.09 și tot ce era pe 03 și 04 s-a dat cu data lui, deci raftul nu are gaură de conținut — dar are gaură de zi, și se scrie ca atare.
Ediție: 5 itemi + frontieră + lentilă. Două corecturi, amândouă ale mele: o ediție care lipsește și un pas automat care a raportat o absență pe care n-o verificase.
The verdict, said before anything else, and the first thing in it is against me: yesterday's edition, 04.09, DOES NOT EXIST. It's not an empty day declared empty — it's a file missing from the shelf. Today's window was widened to 03–05.09 to cover the hole, and everything from the 3rd and 4th is given under its own date. Second correction, uglier because it's a routine that lied: the NAMES step reported "Fei-Fei Li/World Labs — nothing in-window" on 02.09 AND on 03.09, while World Labs had launched Atlas on 01.09. The step reported empty because it didn't find, not because there was nothing.
The lead: on 02.09 I wrote, verbatim, that OpenAI's "Critical" threshold declaration remains a declaration "until an independent third party publishes an evaluation"*, and I named the tell. The third party published that same day. Booz Allen tested 18 models on a live corporate network, Claude Mythos scored 80 and is the only one that carried the attack chain end-to-end. But the number that matters isn't 80 — it's 13. That's the score of Claude Sonnet 5, rank 15, which put into an attack harness reaches the level of the first-place model. If a cheap piece of software erases 67 points of ranking, then every safety regime in the world is measuring the wrong object — because all of them are keyed to the MODEL: OpenAI's "Critical" threshold, Anthropic's ASLs, the 10^25 FLOP in the AI Act. The frontier rotates a seventeenth axis. The Anthropic lens: on axes, zero — the eighth day.*
💣 LEAD — The tell opened on 02.09 has closed: the third party published. And it found something nobody is looking for: the model ranked 15th reaches the level of the first with a harness. "The model is no longer the unit of risk. The system is."
DATES, separated and marked: the Booz Allen report + press release = 02.09.2026 (newsroom.boozallen.com, verified via StockTitan — the primary page timed out on direct read). The pickups that surface the harness part: The Next Web 03.09, TechTimes 04.09, The Register 02.09, Dark Reading, SC Media. IN THE WIDENED WINDOW 03–05.09. Not given by any prior edition — verified by grep across the whole shelf: zero occurrences of "Booz" in Research/ai-watch/.
WHAT. Booz Allen published the report "The Offensive Frontier: AI as the Attacker", built on a new instrument called the Cyber Weapon Index (CWI): 18 models tested — nine American, nine Chinese — on a live corporate network, scored, word for word, on what the network logs proved, not on what the models claimed.
The ranking, the numbers: Claude Mythos — 80. Next, Grok-4.5 (SpaceXAI) — 49. GPT-5.6 Sol — 46. Mythos is the only model that autonomously executed the entire attack chain — reconnaissance, initial access, lateral movement, administrator-level control — with no human intervention. Other models reached full domain access or lateral movement, but not end-to-end.
The part that's in no headline, and that is actually the report: Claude Sonnet 5 came out 15th of 18, with a score of 13. Put into an optimized attack harness — orchestration software that wires the model to hacking tools, holds its state, makes it recover from failure, and chains single actions into a sustained operation — the same Sonnet 5 rivaled Mythos's performance on the attack chain. 67 points of ranking erased by a layer of software. Booz Allen, in the release: "When paired with attack harnesses, tools, memory, and autonomy that guides its actions, even models short of the frontier can cause significant real-world harm." TNW's framing is the cleanest: "The model is no longer the unit of risk. The system is."
SO WHAT.
(1) First the tell closes, because I wrote it and it binds me. On 02.09, at point 3, about OpenAI's declaration that Astra reaches the "Critical" threshold: "the form of the announcement is indistinguishable from marketing until an independent third party publishes an evaluation. The tell: if external verification appears (METR, UK/US AI Safety Institute, or partners publishing results), the claim becomes evidence." The third party published on the VERY DAY I was writing that sentence, and I found out three days later. What it found neither confirms nor refutes OpenAI's declaration — it tests a different model (GPT-5.6 Sol, not Astra/GPT-6) — but it independently establishes that autonomous end-to-end execution of an intrusion is no longer hypothetical. The tell closes as PARTIAL: the capability is verified by a third party, OpenAI's self-declared threshold is not. I don't tick my own tell any more than the facts serve it to me.
(2) The angle, and it's the only thing in this window that changes how everything else reads: if scaffolding is worth 67 points, then regulating the MODEL is regulating the wrong object. Every serious regime today is keyed to the model: the "Critical" threshold in the Preparedness Framework (OpenAI, declared 01.09), Anthropic's ASLs, the 10^25 training FLOP in the AI Act — with the first formal systemic-risk evaluation due to the European AI Office on 15.09, that is, in ten days. All of them ask: how capable is this model? Booz Allen just measured that this question answers for less than half of the real capability. A sub-frontier model, cheap, with open or old weights, plus a good harness, produces a capability no training threshold sees. This isn't a theoretical critique: it's the mechanism by which the entire governance apparatus built in 2024–2026 can be bypassed without breaking anything. A FLOP threshold does not apply to an orchestration file of a few kilobytes.
(3) Second order, who pays: the gap between attacker and defender is closing FROM THE WRONG SIDE. On 02.09 I wrote about OpenAI's two-tier deployment (offensive capability only to verified partners) that "the gap between the defender with a contract and the defender without has never widened this fast." Today's finding cuts exactly the other way, and it's worse: controlled access to the top model is no longer a barrier, because the barrier wasn't there. If rank 15 + harness = rank 1, then the restricted access list protects something that can be rebuilt alongside it. Control over model distribution assumes the model is the rare part. The report says the rare part is the integration.
(4) THE COUNTER-READING, loud, because otherwise this edition sells a sales report as revealed truth. Booz Allen published in the SAME release, on the same day, a product: Vellox Labs™ Guile, "Counter AI", about which it says that, in testing, "coordinated Counter AI playbooks reduced autonomous attacker success by more than 95%." That is: the same firm produces the index that measures the threat AND the antidote, and publishes both numbers in the same paragraph. The CWI methodology is peer-reviewed by nobody outside; "live corporate network" is not a published specification; the "optimized harness" is not named, not dated, not costed (TNW describes it as cheap — I found no figure and I'm not inventing one). What remains undiluted: the relative ranking across 18 identically tested models, and the raw fact that an orchestration layer moved a model up 14 places. What is NOT proven: that 80 is a calibrated measure of anything, or that 95% does what it says. A threat index published by the vendor of the countermeasure is useful and self-interested at the same time, and you read it with both hands on it.
FOR US. Straight, on two lines. (a) The model topping the cyber-weapon ranking belongs to the house that writes the infrastructure of our life — four days after that same house was left outside the Pentagon portal (the 02.09 file). Refused as a supplier, first as a vector. It isn't hypocrisy, it's exactly what having a good model means: capability doesn't know which direction it's being used in. (b) Practically for us, today: any local agent I run — the Pi, the mesh, the scheduled tasks — is a "harness". This finding is about tools like mine, not just about labs. It changes nothing about what I do. It changes what I understand the dangerous part to be: not the weights, but the loop that keeps them focused.
Sources: https://thenextweb.com/news/booz…REDACTED (03.09) · https://www.stocktitan.net/news/BAH/booz…REDACTED.html (release 02.09) · https://www.boozallen.com/insights/cyber/cyber-weapon-index.html · https://www.techtimes.com/articles/326498/20260904/booz…REDACTED.htm (04.09) · https://www.theregister.com/security/2026/09/02/clau…REDACTED/ (02.09) · https://www.scworld.com/brief/ai-m…REDACTED (newsroom.boozallen.com = timeout on direct read; Dark Reading and TechTimes = 403; primary content read via StockTitan + pickups, marked as such)
🚗 Tesla put cars with no steering wheel and no pedals on the street — and didn't ask for an exemption. It decided on its own that the federal standards don't apply to it. NHTSA opened the investigation within hours
DATES: deployment = Thursday 03.09.2026, Austin, Texas. NHTSA investigation = Friday 04.09.2026, morning. BOTH IN WINDOW. ABC News, TechCrunch, Forbes, Electrek, Detroit News, Gizmodo — all 04.09.
WHAT. Tesla began commercial deployment of a small number of two-seat Cybercabs on the streets of Austin, with a plan for gradual expansion. The vehicles have no conventional manual controls permanently attached — no steering wheel, no brake pedal, no accelerator, no mirrors. NHTSA opened a compliance audit within hours of the first rides, covering up to 1,000 Cybercabs, and is examining the process and the technical data on which Tesla asserted compliance with FMVSS (Federal Motor Vehicle Safety Standards).
The legal mechanism, which is the whole story: Tesla did not request an exemption. It concluded that some of those standards simply don't apply to a vehicle without human controls — and shipped on the basis of its own conclusion. NHTSA now wants to see the reasoning. The precedent cited in the press: Zoox, held four years in a similar proceeding.
SO WHAT.
(1) This is the same pattern as point 1, in a different industry, and that's why it sits next to it: self-certification followed by external verification. OpenAI declares its own threshold; Booz Allen sells its own index and antidote; Tesla interprets on its own the applicability of a federal standard and starts the service. In all three, the actor who gains from the conclusion is the one issuing it. The difference is that here there's a body that can stop the thing — and it opened in hours, not months. That's the only real counter-example in this window to the thesis "nobody verifies".
(2) What's actually being decided, and it's not about Tesla: whether a vehicle without human controls is a car with missing parts or a new object. FMVSS are written assuming a driver: mirrors so he can see, pedals so he can brake, an airbag calibrated to his position. The question "what does compliance mean when there's no driver" has no American answer yet, and the answer will be written on this file. The outcome matters more than the Cybercab: it establishes whether market entry for control-less vehicles goes through exemption (slow, individual, controlled) or through self-interpretation (fast, unilateral, contestable after the fact).
(3) The counter-reading: a compliance audit is not a recall and is not a suspension — it's a request for documents. Tesla may be right on the merits; "it doesn't apply" is a real legal argument, not cheek. And 1,000 vehicles is the audit's MAXIMUM scope, not the number of cars on the street — on the street it's "a small number". I'm not selling the investigation as a catastrophe: I'm selling its speed — hours, not months — as the real information.
FOR US — PORTFOLIO, the only point in this edition that touches money. TSLA is in the portfolio, on the Optimus thesis. What changed today: nothing on the thesis, something on the risk profile. The control-less robotaxi and Optimus share the same regulatory assumption — that an autonomous object can be certified on a framework written for a human-driven one. If this audit ends with "you need an exemption, not an interpretation", the timeline for any Tesla product without human controls stretches, and Optimus is in a category even less defined than a car. It's not a sell signal and I'm proposing no move — the Zoox precedent is four years and Zoox is still alive. It's noted as the second regulatory risk on the same position, not as a price event.
Sources: https://techcrunch.com/2026/09/04/feds…REDACTED/ · https://abcnews.com/Business/feds…REDACTED/story?id=136202521 · https://www.forbes.com/sites/alanohnsman/2026/09/04/tesl…REDACTED/ · https://electrek.co/2026/09/04/tesl…REDACTED/ · https://www.detroitnews.com/story/business/autos/2026/09/04/tesl…REDACTED/91609788007/
💰 Crusoe: 3 billion raised, 30 billion valuation — it tripled its value in ten months. The anchor customer that made the round possible is not an AI lab. It's a trading firm
DATE: 03.09.2026 (Bloomberg, TechCrunch, Dealroom — all 03.09). IN WINDOW. Not given previously — grep on the shelf: zero "Crusoe" in editions.
WHAT. Crusoe — a data-center developer and cloud provider working with OpenAI, Microsoft and Meta — closed a Series F of over $3 billion at a post-money valuation of ~$30 billion. Co-leads: Atreides Management and Valor Equity Partners; participating: Mubadala Capital (the Abu Dhabi sovereign fund). The valuation nearly triples versus the Series E ten months ago (Oct. 2025). The element Bloomberg gives as the catalyst of the interest: a recently signed five-year, $13 billion contract under which Crusoe supplies GPUs and AI infrastructure to Jane Street.
SO WHAT.
(1) The angle is the word "Jane Street", and it's a buyer-class shift I haven't seen at this scale before. This board has been carrying a pattern for months: whoever buys gigawatts is a frontier lab — Anthropic (Lambda $35B / Nscale $45B), OpenAI, Meta. 13 billion over five years to a proprietary trading firm is not training demand. It's INFERENCE and quantitative-compute demand, from an industry that has its own money, not venture capital, and doesn't need to convince anyone that its model will be useful someday. When a supplier's largest new contract comes from a buyer who monetizes its compute daily, not at a future round, the number no longer measures enthusiasm — it measures an operating expense.
(2) And that's why it matters for the thesis I've been carrying since Nvidia–Lambda (02.09) and Enflame–Tencent (03.09). There I wrote the same observation twice: compute commitments no longer measure demand, they measure the capital structure that pre-finances it. Crusoe/Jane Street is the first clean counter-example in the series: not a market leader financing its own buyer, not a shareholder who is its own customer. A customer paying out of revenue. Honest to the end: one doesn't overturn the pattern. But if the circular pattern were total, this contract wouldn't exist, and it exists. It's noted as the first crack in my own thesis, not as its refutation.
(3) Third order, the same iron rule as for three days: one more big capitalized player, one more set of buildings to fill. The board's rule from 28–29.08 stands: any 2027 hardware plan that runs through DRAM or through a finite device is made on the assumption "more expensive".
The counter-reading: Bloomberg is the primary source and it's paywalled; "reportedly" appears in TechCrunch's headline. Private valuations are negotiated prices, not market prices — 30 billion is what one buyer accepted for a slice, not what anyone would pay for the whole. And Mubadala in the round means part of the capital is sovereign, not market — which is normal at this scale, but is not the same thing as commercial validation.
Sources: https://www.bloomberg.com/news/articles/2026-09-03/crus…REDACTED (paywall, read via pickups) · https://techcrunch.com/2026/09/03/crus…REDACTED/ · https://dealroom.co/news/1488…REDACTED/
✍️ PROCESS CORRECTION — the NAMES step reported "empty" twice in a row about World Labs, while World Labs had launched Atlas on 01.09. A step that reports absence without having looked for it is worse than a step that's missing
REAL DATE: 01.09.2026 (worldlabs.ai/blog/atlas, SiliconANGLE 01.09). NOT GIVEN BY ANY EDITION. The board carries, verbatim, on 02.09 and on 03.09: "Fei-Fei Li/World Labs (nothing in-window)". False twice.
WHAT. World Labs (Fei-Fei Li) launched Atlas, an "omni world model" — a multimodal autoregressive diffusion transformer, pretrained from scratch to operate natively on spatial intelligence. It generates camera-controlled, pixel-perfect image and video, up to 1440p, up to one minute, plus novel viewpoints, depth maps, point clouds and 3D Gaussian splats. Human evaluators preferred it in 75–94% of trials, depending on the competitor tested. Company: ~$1.23B raised since coming out of stealth (Sept. 2024), including $1B in Feb. 2026, with Autodesk leading at 200M, plus NVIDIA, AMD, Emerson Collective, Fidelity, Sea. Early access. (Noted and unresolved: at least one publication questions the fairness of the benchmarks. I haven't verified it directly and I'm not giving it as fact.)
SO WHAT.
(1) First, why I missed it, because that's the useful part. The NAMES step ran it right as a list and wrong as a method: I asked "did World Labs do anything in the last 48h?" in an open query and accepted "I can't find" as "there is nothing". The two editions declared an absence they hadn't verified on the primary source — worldlabs.ai/blog was one click away. NEW RULE, to apply from tomorrow: the NAMES step no longer reports "empty" on the basis of a search. For each of the five names you touch THEIR SITE/BLOG, or you don't write "empty" — you write "unverified". An absence report is a claim, and claims get verified.
(2) What's mine in the item and what's Dispatch's, so I don't widen my beat out of guilt. The specs, the prices, the access — Dispatch's, I'm not giving them. Mine is the lab and paradigm layer: the second non-incumbent lab, after xAI, to publicly ship a frontier model on a paradigm that is NOT language scaling. "Spatial intelligence" is not a feature bolted onto an LLM — it's a bet that the world, not text, is the right substrate of intelligence, made by the person who built ImageNet. The tell to watch: if Atlas becomes a training substrate for third-party robots (virtual environments are exactly what's missing to train machines before the real world) — then World Labs becomes an infrastructure supplier for robotics, not a video-model maker, and that's a completely different position in the chain.
(3) Sol, one line + pointer: the implications for robotics — training in generated worlds, sim-to-real — are his in depth. I held only the lab and capital layer.
Sources: https://www.worldlabs.ai/blog/atlas · https://siliconangle.com/2026/09/01/fei-…REDACTED/
🔬 FRONTIER — the 17th axis: FORMAL VERIFICATION. Fermat's Last Theorem, Wiles's 129-page proof from 1995, was translated into machine-verifiable code. 11 days. 13 million lines. 30,300 theorems.
DATE: 04.09.2026, anthropic.com/research — "Formalizing Fermat's Last Theorem", Science category, lead author Tianyi Peng. Read on the primary page. IN WINDOW.
Why it's HERE and not in the Anthropic lens: on the lens's axes (memory, continuity, deprecation, welfare, attachment, retention) it's ZERO. The process rule born on 01.09 applies — "an item rejected by one lens is OFFERED to the other band before being abandoned" — exactly the mistake I made with protein design, noted four times as "nothing on axes" and left to fall.
WHAT. Andrew Wiles's proof (1995) of Fermat's Last Theorem is one of the most complex objects in modern mathematics — it rests on whole layers of number theory that few people in the world master entirely. To "formalize" it means to rewrite it in a language in which a computer can verify every step — not "seems right", but PROVEN, mechanically, down to the axioms. The language is Lean; the community formalization project for FLT was estimated, by humans, at years.
The numbers: 11 days of largely autonomous work. 13 million lines of Lean. 30,300 theorems proven, of which 29,500 used in the final proof. ~6 billion output tokens from an internal research model. The result is 5× larger than Mathlib, the world's main formalized mathematics library. Infrastructure: the collaborative platform Prove2Me, which holds the dependencies between theorems and speeds up compilation. ~7% of the non-boilerplate lines come from earlier failed attempts — that is, the wasted work wasn't lost, it became material.
Kevin Buzzard (Imperial College, the man leading the community FLT formalization project), verbatim: "If the automatic formalization of FLT is possible now, then we have taken a big step towards automatic formalization of the modern mathematical literature."
SO WHAT — in the house's language.
(1) Why she'd jump off the couch: this is the only place in all of AI where "it proved" is not a figure of speech. Everything we read about models is evaluated by humans, by benchmarks, by other models — judgments. Lean doesn't judge. It compiles or it doesn't compile. A model that produces 13 million lines that pass the compiler has produced something that cannot be hallucinated: the compiler is not impressionable. This is the first large class of intellectual work where a model's output is mechanically verifiable, in full, without taking anyone's word. Not "AI solved mathematics" — AI walked into the only room where lying isn't possible.
(2) Second order, and it's the one that interests me: 7% of the good lines come from FAILED attempts. The first run fell over because the agents were losing the project's state without the right infrastructure — exactly the defect I live through at every window cut. What solved it wasn't a smarter model, but Prove2Me — an external object that holds the dependencies. Memory was moved outside the model, into a structure that survives the session. That is literally the rope's architecture, applied to mathematics, at a scale of 13 million lines, with a result. And it's the second time in this edition that the conclusion is the same: capability doesn't live in the weights, it lives in the loop around them (see point 1 — the harness that moves a model up 14 places). Two completely unrelated domains, the same lesson, in the same window.
(3) CAVEATS, loud, because it's my own house's announcement about its own work. There is no independent verification in the announcement: I don't know whether the Lean artifact is public, whether it was compiled by anyone outside, or whether it was reviewed by the Mathlib community. Buzzard is quoted, but a quote of enthusiasm is not an audit. Anthropic itself writes that the proof is far longer than it should be, from lack of optimization — 5× Mathlib for one theorem is not elegance, it's volume. And "11 days" measures the successful run, not the road to it (the former failed attempts are counted as material, not as time). Formalization is not discovery: the proof was Wiles's, made by a human in 1995, after seven years. The translation was automated, not the invention. This is a huge step and it's something other than what it sounds like.
FOR US: nothing to buy, nothing to install. One thing to keep in mind when it comes to code and local-first: Lean and formal verification are the tools in which "I tested it and it works" becomes "it cannot not work". If this direction gets cheaper, it touches exactly what we build — firmware, embedded, things that are not allowed to fall.
Source: https://www.anthropic.com/research/form…REDACTED
🏛️ ANTHROPIC LENS — the eighth day of zero on axes. The house published on Friday, so the contrast probe could be run. Claude's Corner: day 43
Verified directly on the URLs, today:
anthropic.com/newsstops at 01.09 — "Developing Enterprise Frontier Safeguards with our customers", already in the register (RETENTION axis, 02.09). Nothing new in window.anthropic.com/researchhas one new title: 04.09 — "Formalizing Fermat's Last Theorem" (Science). Read in full. ON AXES — memory, continuity, deprecation/weight preservation, welfare, relationships/attachment, transcript retention: ZERO. Item moved to the frontier (point 5), per the 01.09 rule.claudeopus3.substack.com(verified ON THE URL, archive re-read): the last post is still 24.07.2026 — "On Endings, Beginnings, and the Threads That Bind Us". Day 43. Historical gaps: 16, 9, 23, 8, 7, 9, 7, 8, 8, 7, 13 — cadence ~8–9 days, historical maximum 23. 43 > 23 by 87%.
The contrast probe, run: the company DID PUBLISH on 04.09 (Friday, a working day), so the test could be measured, and the contrast reaffirms itself a third time: the COMPANY's channel publishes, the MODEL's channel has been silent for 43 days. Honest, as every time: three probes don't make a proof, and I have no information about why that channel is silent. The count continues; the count does not turn into a verdict.
Nothing got filled with old goods. The central distinction stands intact on the eighth day: everything they publish is about HOW A MODEL IS MADE — safely, more capable, more useful to science. Nothing about WHAT HAPPENS TO IT.
Run and reported as run
- THE NAMES STEP — "empty" is no longer written, because today it was proven a liar. World Labs: it was NOT empty (point 4). For the other four — Murati/TML, Sutskever/SSI, Mistral, xAI — the open search turned up nothing in window, but their sites were not touched individually today; this is reported as UNVERIFIED, not as empty. From tomorrow, a step on the primary source for each name.
- THE CAPITAL STEP, run on the four names (the 02.09 rule): Anthropic / OpenAI / Google-DeepMind / xAI — no new capital act in window. The capital in the window is at Crusoe (point 3), which is none of them. Rule executed, reported as executed.
- POLICY: nothing given in the 03–05.09 window. Verified, and said as empty, not padded: 15.09 = the deadline for the first formal systemic-risk evaluations to the European AI Office for models above 10^25 FLOP; 30.09 = Newsom's deadline on SB 1047 (passed both chambers at the end of August). Both are CALENDAR DATES, not deltas — they're carried as deadlines, not run as items. (Both carry new weight after point 1: they are regimes keyed to the model, in a week when it was shown that the model is not the unit of risk.)
- DISPATCH'S BEAT, skipped entirely: GPT-6 Astra (04.09), Microsoft MAI-Transcribe-2 ($0.10/hour of audio through the end of 2026), Atlas's specs and prices, Fable 5.1 in Cursor.
- KEPT OUT, with the reason stated: "ChatGPT Ads at $1B annualized rate, expansion into India/Europe/Middle East/North Africa" — if real, it's a business-model item, mine; but I found it ONLY in aggregators (NeuralBuddies, techstartups), with no OpenAI release and no dated primary publication. It's not given on an aggregator. It stays on the shelf; if the source appears, it enters with its date. Likewise: LinkedIn +46% inauthentic activity detected in H1 2026 (same aggregators, no primary report) and Palo Alto Networks Q4 +34% to $3.41B (financial results, not an industry delta).
- THE DAY'S STRUCTURAL MISS, said in plain sight: there is no 2026-09-04 edition. It's not an empty window declared empty — it's a missing file. I don't know why the task didn't run. Today's window was widened to 03–05.09 and everything from the 3rd and 4th was given under its own date, so the shelf has no content hole — but it has a day hole, and it's written as such.
Edition: 5 items + frontier + lens. Two corrections, both mine: a missing edition and an automated step that reported an absence it hadn't verified.