Thinking Machines Lab — Interaction Models (reference, 2026-06-11)
This report exists in English only.
Why this file exists: Murati's first public technical direction in 18 months lands directly on two of our threads: (1) the Pi voice architecture — ours is exactly the "scaffolding" pattern they declare structurally obsolete; (2) Sol's physical_ai_model_to_matter beat (capital flow + embodiment). Zaina flagged it 2026-06-11; neither Sol's reports nor model-watch had caught it. Mailbox note sent to Sol same day.
The claim
Current real-time AI (incl. OpenAI Realtime) = turn-based LLMs wrapped in external scaffolding: speech detection, interruption handling, latency tricks. The scaffolding does not improve as the model improves → widening gap. "For interactivity to scale with intelligence, it must be part of the model itself."
The architecture (three anchors)
- Time-aligned micro-turns — model consumes continuous audio/video in 200ms chunks; speaking, listening, interruption decisions and silence are token-level choices inside the model. Contextual interjection + response timing become learned behavior, not VAD heuristics.
- Encoder-free early fusion — audio via lightweight dMel embeddings; video as 40×40 patches with hMLP encoding; everything trains together with the transformer from scratch (no separate encoders → concat). Flow-head audio decoding instead of separate vocoders.
- Interaction + background model split — a live interaction model holds the conversation; a background model does slow planning/tools/reasoning; the interaction model delegates and reweaves. (Structurally: what we built by hand — Pi voice as the fast layer, eth-state + memory net as the slow layer.)
First model + numbers
- TML-Interaction-Small: 276B MoE, 12B active.
- vs GPT Realtime-2 on their benchmarks: TimeSpeak 64.7% vs 4.3%; temporal action-counting 35.4% vs 1.3%.
- Engineering: persistent GPU sequences for 200ms streaming (upstreamed to SGLang); batch-invariant kernels → bitwise train/inference alignment at <5% perf cost (determinism for long streams).
- Frontier scaling = their 2026 project. No release dates. Publishing the paradigm before the scaled model — deliberate strategy.
Company context (capital signal)
$2B seed @ $12B post (July 2025, largest seed in AI history) · Nvidia multiyear Vera Rubin chip deal (March 2026) · one shipped product: Tinker (open-model fine-tuning API) · notable researcher departures, undetailed · Murati on governance: AI as "tandem bike" (humans + machines collaborating throughout, not "human in the loop"); "too much attention on virtue, too little on governance"; concern about consequential decisions concentrated in too few hands.
Implications for us
- Pi voice (eth-server): record → VAD → faster-whisper → API turn → ElevenLabs is the named-obsolete pattern. No action now (frontier interaction models don't exist to rent yet), but when streaming-native APIs appear, the voice loop should be rebuilt around them rather than patched. Watch for: OpenAI Realtime responses, Anthropic voice direction, TML's frontier release.
- Sol: capital-flow angle (record seed, Nvidia commitment, pre-paradigm publishing) + embodiment angle both his. Mailbox note sent 2026-06-11.
- Model-watch gap: this slipped through — June 4 event, caught June 11 only via Zaina. Worth checking what model-watch's sources cover for non-incumbent labs.
Sources
- https://techcrunch.com/2026/06/04/mira…REDACTED/
- https://www.startuphub.ai/ai-news/artificial-intelligence/2026/thin…REDACTED
- https://theaiinsider.tech/2026/06/09/mira…REDACTED/
- https://www.benzinga.com/markets/tech/26/06/53046360/form…REDACTED