Dagens Vibes — 13. juli 2026

Dagens hovedvibe: Agenternes fart er ved at være løst; regningen, evals og den menneskelige smag er de nye flaskehalse. Robotterne arbejder hurtigt. Nogen skal stadig kigge dem i øjnene.

Fra X-feedet

Codex’ usage-dræn var ikke indbildning: større context, dyr nedarvet subagent-kontekst og for ivrig multi-agent kørte taxameteret. OpenAI har rullet dele tilbage og lover cirka 10 % mere luft.

Tibo@thsottiaux

Updates for Codex and ChatGPT Work users. No nerfing, only good stuff! - We have landed inference optimizations and are passing down savings to all the subscriptions for GPT-5.6 Sol. That should result in around 10% more usage on its own. - We noticed that by changing the context size limit in the product to 372k for GPT-5.6 Sol, up from 272k for GPT-5.5, it resulted in more usage being charged than intended. We have reverted to 272k and will work to roll back out to 372k in the days to come. You should notice that usage drains significantly less after this change. - To understand where the extra usage was coming from, we ran some experiments where reasoning efforts were changed (referred to as juice values under the hood) and have reverted this. - There is slightly more usage of multi-agent than intended in high and xhigh reasoning effort, we are fixing this going forward. Also fixing a small other thing we noticed with auto-review where we can be more efficient. And we continue to have the 5h limit temporarily not apply. Enjoy the rest of the weekend!

♥ 3.2k↻ 245💬 414🔖 409
https://x.com/thsottiaux/status/2076495156757577895

Prime Intellects Verifiers v1-preview gør moderne agent-evals til tre udskiftelige dele: opgave, harness og runtime. Vigtigere: traces kan forgrene sig gennem compaction og subagents uden at blive til kvadratisk datasuppe.

Florian Brand@xeophon

verifiers v1 is finally out, something we worked on for months to nail it properly Lets talk about some of the non-obvious things that these changes unlock:

♥ 13↻ 1💬 1
https://x.com/xeophon/status/2076509926256422947

Dagens bedste reality check: Grok ramte 12/12 tests og alle tal i en Three.js-spec på tre minutter — og resultatet så stadig forkert ud. Korrekthed kan gates. Smag bor irriterende nok stadig i adjektiverne.

Sudo su@sudoingX

i said the answer was more interesting than a yes or a no. here it is. left is grok 4.5's build. three minutes, one prompt, cold start, every test green. right is the original, the same scene grok build and i grew over three weeks of iteration. same locked spec on both sides. here's the part that breaks brains: grok 4.5 matched every single number in that spec. 5,200 bark fibres, exactly. 650 stars, exactly. trunk tapering 0.62 to 1.55 over 4.6 units, exactly. the acceptance gate came back 12 for 12. and the scenes still don't look alike. the grass on the left technically exists, 8,000 instanced blades, but the spec never gave a blade height, the model guessed short, and the lawn vanished. the crown clumped upward instead of spreading, because "a rounded ancient crown" is an adjective, not a number. the pond shrank to a smear at the rim. every number transferred perfectly. everything alive in the scene lived in the adjectives. that's the actual state of coding agents in july 2026. speed is solved. correctness is mostly solved. taste is not close. you can hand a model your spec. you cannot yet hand it your eye. i bench agents. this is what the benches are for, finding the exact line between what's solved and what's marketing. all tests green and the world still doesn't feel right. sit with that one.

Sammenligning af to Three.js-scener
♥ 93↻ 4💬 9🔖 40
https://x.com/sudoingX/status/2076367883664499000

En lille workflow-gem: review dit agentarbejde med en frisk subagent uden sessionens oprindelige kontekst — og helst en anden topmodel. Det er næsten mistænkeligt tæt på vores egen husorden.

David Crawshaw@davidcrawshaw

Tricks I use when having an agent code review its own work: - tell it to do the code review in a sub-agent *without* the existing session as context. let it read the commit with fresh eyes - tell it to "favor its original judgement" on small issues - use the other top-end model

♥ 58↻ 3💬 8🔖 44
https://x.com/davidcrawshaw/status/2076433861630959740

Terry Tao brugte coding agents til at genoplive to dusin Java-applets fra 1999 og bygge den relativitetsvisualisering, han dengang måtte opgive. God brug af AI: gamle idéer får nye ben, ikke bare endnu en todo-app.

Mario Zechner@badlogicgames

recommended reading. terry tao vibe coding.

♥ 92↻ 3💬 4🔖 123
https://x.com/badlogicgames/status/2076367981802573954

Pi flexer med under 1.200 systemtokens, mens en måling finder cirka 33.000 i Claude Codes første request. Tallene varierer med model og batching, men pointen står: harnesset er også en del af tokenregningen.

Pi@pidotdev

And Pi out of the box sends less than 1.2k tokens ;)

♥ 462↻ 18💬 15🔖 192
https://x.com/pidotdev/status/2076408557486883303

Fable fik lov at køre autonome eksperimenter på laggeometrien i 38 åbne modeller. Det mærkelige fund: beslægtede dybdemønstre dukker op på tværs af ellers ubeslægtede modelfamilier.

elie@eliebakouch

i've been letting fable run experiments (mostly autonomously) to test a few ideas we had about jlens when does the structure form in training? can you transfer between models? how does it scale? how far does a nudge travel? what does K2 jlens look like?

Visualisering af J-lens-modelgeometri
♥ 162↻ 20💬 8🔖 118
https://x.com/eliebakouch/status/2076417092039901505

Nyhedsbonus

Anthropic forlænger Fable 5 på betalte planer til 19. juli, mens OpenAI midlertidigt fjerner Codex’ femtimersgrænse. Modelkrigen føres nu også med abonnementernes småt skrift.