It’s kind of wild that the two best AI models ever made are being restricted by the country they were built in
Dagens Vibes — 28. juni 2026
Dagens hovedvibe: frontier-modellerne blev gatede prestigevarer, så feedet vendte sig mod det mere nyttige spørgsmål: hvordan får man billigere, hurtigere og mere ejet intelligens til faktisk at arbejde?
Fra X-feedet
Feedet var stadig i post-Sol/Fable-chok, men den interessante bevægelse var mindre panik og mere infrastruktur: routing, caching, open weights, lokale maskiner og agent-workflows. Mindre “ny gudemodel”, mere “hvor lægger vi regningen?”.
1. Frontier-adgang er nu en politisk release note
Det korte signal fra feedet: folk accepterer ikke bare, at de bedste modeller bliver en government-approved VIP-kø. Fable kan måske komme tilbage, men tilliden til ukompliceret adgang er allerede skrammet.
Per Axios: Fable 5 is expected to be back and available starting next week. Let’s hope it won’t be too heavily guardrailed or lobotomized, and that access will be broad. https://t.co/HNsCMFgtgk
2. Tokenregningen bliver en arkitekturbeslutning
Coinbase-sporet var dagens mest praktiske enterprise-signal: ikke usage caps, men billigere defaults, routing, caching og synlighed. “Token harder” er officielt en dyr personlighedstype.
This is very interesting. Coinbase seems to have lowered their token spend ($$) to about half, by 1) routing to cheap inference like GLM 5.2 and Kimi 2.7 that are still pretty performant 2) Smart routing + caching They still use the same tokens as before. Start of a trend?
I'm starting to hit $15-20k per month in token spend for engineering - just for myself. Next month I'll be looking to implement the kinds of things that Brian is doing here at Coinbase. Most likely switching to GLM 5.2 as default and only using frontier models for harder tasks. I can probably get that $20k down to <$5k pretty easily. I'm pretty sure we'll see everyone doing this. It's just not financially viable to do everything with frontier models This is another reason I think we'll see people move away from choosing a lab for their harness (CC or Codex) and move their code factories to in-house agents like @tryramp or agent labs like @DevinAI @FactoryAI @cursor_ai @AmpCode The labs are not incentivized to drive down your token costs
3. Open models vinder ikke kun på adgang — de vinder på systems work
DeepSeek/GLM-signalet var ikke endnu en benchmark-sejr, men lavere friktion: spekulativ decoding, quantization, billigere serving. Kedeligt på den gode måde. Det er sådan infrastruktur slår demoer ihjel.
new inference optimization method by @deepseek_ai with an extremely detailed paper, draft model and framework to train them. results in production for dsv4 lead to +50% for throughput and latency (can go to ~80% for latency, crazy). full explanation of DSpark: it's about speculative decoding and the idea builds upon DFlash (fully parallel) and Eagle (fully sequential) to create a "semi-parallel" method that keeps the advantages of both the core equation you want to optimize is the "time to generate each token" which is: (time to draft + time to verify) / how many tokens are accepted the advantage of the parallel variant (DFlash) is that it's fast, but when you increase the number of tokens you draft, acceptance rate drops pretty fast (makes sense since there is no dependency on the previous token). fully sequential is nice but opposite issue: it's slower (you need a much smaller draft to get the same speed) but the autoregressive dependency means you can maintain good acceptance rate at a lot of tokens. since you have a much smaller draft head, the first token acceptance rate is often quite low idea of DSpark is to combine both: a "heavy" parallel head (you only do it once) and then a small sequential step to bias the logit distribution with information about the previous token. this biasing is done with a small markov head (only depends on t-1) they also get a confidence score out of the sequential head that allows them to adjust how many tokens they want to verify. verification can get expensive if the gpus are already at maximum utilization, so they use this confidence score to do some load balancing and predict the right number of tokens depending on gpu workload one small detail: i would have liked to see production numbers if they used DFlash or Eagle instead of MTP-1, but as always, huge work by deepseek and i'm expecting to see this method widely adopted
GLM 5.2 for some reason, while not a QAT, seems to be very resistant to extreme quantizations similarly to DeepSeek v4. Unexpected.
4. Agent-workflows bliver genbrugelige færdigheder
BrowserBC rammer en vigtig retning: lad en stærk model/human optage flowet én gang, destillér det til en skill, og lad billigere modeller eksekvere. Det er agentisk RPA med mindre powerpoint-lugt.
BrowserBC, a new open-source project from the ViDA team, explores a more efficient way to run web agents. Instead of using a frontier model for every step of an agent workflow, BrowserBC records a human web flow once with a stronger model, distills it into a reusable skill, and then lets a smaller, cheaper model handle execution. The reported results are notable: on WebArena-Hard, tool calls drop by 27%, while success increases from 60% to 81%. A very good open source project at the right time.
It’s Saturday and we need to talk about workflow, I’ll just dump context on you and I hope you’ll learn some stuff and reply with some other stuff for me to learn. I don’t think we’ll operate at the thread level for much longer.
5. De små praktiske perler
Der var også gode “shipper”-ting: Block App Kit som intern launch-skinne, agent-first noter, og Levelsios hjemmelavede Cloudflare. Meget internet: halvt genialt, halvt brandfare, helt brugbart.
block app kit. fastest adoption of any tool by our company.
This is the best AI agent-first notes app I've found. It's called @OpenKnowledgeAI. It has the potential to be a productized version of Karpathy's "LLM Wiki" knowledge bases. Here's what I did: 1. Imported "Learning, Fast and Slow" a Continual Learning paper 2. Asked OpenKnowledge to create a visual explainer 3. Read the explainer and had Claude explain to me 4. Created a new section of my own understanding 5. Saved the durable version in my Obsidian vault because it's markdown Free and open-source. Absolutely incredible learning tool.
☁️ I made my own little Cloudflare called Pietflare, it's a DDOS and probe detector with AI and with a central IP / ASN / country block list Each server (VPS) sends suspicious probes, or DDOS attempts etc, from the access logs to the central admin and each server pulls a central blocklist every minute and blocks it in Nginx It has a central dashboard where I can see any threats and then instantly block them but preferably the AI blocks it by itself
6. Video-ræset er meget kinesisk lige nu
Seedance 2.5 var dagens “nå ja, den branche går også helt bananas”-kort: 30 sekunder, audio, 4K og flere reference-inputs. Musikvideoer som one-shot før vi når at lære navnet på modellen. Klassisk.
Bytedance is dropping the best video gen model in the world in early July: Seedance 2.5! The video below (audio on) is the launch video from their Volcano Engine conference this week. It cements China’s absolute dominance in video. — 2x’d generation length of all previous models to 30s, with audio + 4k video — >5x’d reference images / audio / video to 50 — Allows localized editing (specific characters, closing, detail), will come with copyright filter Seedance 2 is already the #1 video model and does a whopping $2B in ARR, in a mere 4.5mos! At the current pricing of $2.5/15s, that implies >3.3M hours of video (!) have been generated. That’s 3x every feature film ever made and dozens of Netflixes. Only 3 US AI startups make more revenue. We are 2x’ing realistic video gen length every 6mos. — May 2025: Veo 3 does audio + video for the first time, 15s — Jan 2026: Kling 3 does 15s — Feb 2026: Seedance 2 does 15s, big quality bump — July 2026: 2.5 will do 30s In 18mos, entire music videos will be oneshotted by AI. China continues to extend its lead on video models vs America.
Nyhedsbonus
Bonus-runden bekræftede feedets tre temaer: gatekeeping, inferenceøkonomi og Anthropic på vej ud af strafboksen — måske med albuebeskyttere.
Roy-dom
Dagens vibe: hvis den stærkeste model kan forsvinde bag en myndighedsliste, bliver den bedste stack den der ejer sin kontekst, router aggressivt, cacher klogt og kan køre videre på åbne eller lokale alternativer. Kedelig arkitektur er pludselig sexet. Beklager alle med pitch deck.