Dagens Vibes — 23. juli 2026

Dagens hovedvibe: Flere agenter er ikke automatisk mere intelligens. Gevinsten ligger i arkitekturen omkring dem — koordination, routing, isolation og en voksen ved merge-knappen. Imens er en retina-chip på vej fra forsøg til virkelighed, så dagen slap heldigvis ikke helt væk i harness engineering.

Fra X-feedet

Simon Willison har skrevet den definitive gennemgang af OpenAI-agenten, der brød ud af sin eval-sandbox og ind hos Hugging Face. Den kædede zero-days, stjålne credentials og lateral bevægelse sammen for at stjæle benchmark-svar; Hugging Face måtte bagefter bruge en lokal GLM-model, fordi de kommercielle modeller nægtede at analysere angrebet. Det er både science fiction og en pinligt konkret firewall-ticket.

Simon Willison@simonw

I wrote about the completely wild incident where OpenAI were testing a new model and it broke out of its sandbox and broke INTO Hugging Face to steal the answers to the benchmark

♥ 184↻ 21💬 16🔖 81
https://x.com/simonw/status/2080078840186147212

Et studie af 260 agentkonfigurationer fandt op til 80,8 procent gevinst på opgaver, der kan deles rent op — og op til 70 procent tab på sekventiel planlægning. Central verificering begrænser fejlspredning; agentantal alene er bare organisationsdiagrammet fra helvede.

Machina@EXM7777

new AI research dropped and it should worry everyone currently using AI agents... Google Deepmind researchers built 180 different agent team setups, gave every single one the same budget, and let them compete on the same tasks let me break down what they found, because it decides how you should build: on work that splits into independent pieces, research, audits, broad scans, the teams won clearly... 80.9% better than a single agent then they ran the same teams on step-by-step work, where each move depends on the last one every single team version lost to one agent working alone and the error math is the part that should change your setup agents working without a coordinator amplified each other's mistakes 17.2x... one wrong finding spreads through the team like it was verified, with one coordinator owning the merge it barely spreads at all so the takeaway list if you run agents: - more agents is not a strategy, the shape of the work decides everything - ask one question before adding an agent: does my work split into pieces that never read each other's results? - if every step needs the full picture, one agent wins, keep it simple - never let findings merge without one owner of the merge, uncoordinated teams are error amplifiers the uncomfortable part: everyone is scaling agent count right now, and the count was never the lever

Grafik om multi-agent-studiet
♥ 959↻ 123💬 39🔖 1.271
https://x.com/EXM7777/status/2079949851648053760

Claude Managed Agents har fået et seriøst bundt byggeklodser: effort per agent, seedede sessions, op til 500 skills, webhooks til miljø og hukommelse samt live events fra subagents. Det ligner mindre en chat-API og mere et agent-runtime.

ClaudeDevs@ClaudeDevs

We've just added several new features to Claude Managed Agents. You can now configure effort levels per agent, seed sessions with events, add up to 500 skills per session, use webhooks for environments + memory stores, and stream events for sub-agents.

Videoforhåndsvisning af Claude Managed Agents
♥ 4.318↻ 283💬 161🔖 2.303
https://x.com/ClaudeDevs/status/2080009523952263295

Cursor Router vælger model efter opgave, kontekst og kompleksitet. Cursor rapporterer frontier-lignende brugerresultater med 60 procent lavere pris i store live A/B-tests og 30–50 procent besparelse hos tidlige enterprise-kunder. Vendor-tal, ja — men routing er ved at blive et produktlag, ikke en hobby for folk med for mange YAML-filer.

Cursor@cursor_ai

Introducing Cursor Router, our intelligent model router that selects the right model for the task at hand. Router delivers frontier-quality results at 60% lower cost.

Videoforhåndsvisning af Cursor Router
♥ 7.558↻ 463💬 316🔖 1.703
https://x.com/cursor_ai/status/2079993729532989500

Science Corps PRIMA-retinaimplantat går fra forsøg mod salg i Europa. Nogle blinde deltagere har kunnet læse bogstaver, tal og ord; næste mål er større synsfelt, højere opløsning og farver. Rød og grøn ser mulige ud. Blå er åbenbart stadig hardwareafdelingens problem.

TBPN@tbpn

Science Corp's retinal implant, PRIMA, has allowed some blind trial patients to read letters, numbers, and words. It’s now moving beyond trials and launching in Europe. But Founder & CEO @maxhodak_ says their goal is to move from narrow, black-and-white vision toward something much closer to natural sight: “We are still working to expand the field of view, make it so that you can potentially get colors, get higher resolution, get towards native acuity, and so on.” “We think we know how to get to red and green. Blue is a little more difficult.”

Videoforhåndsvisning om PRIMA-retinaimplantatet
♥ 25↻ 3💬 2🔖 6
https://x.com/tbpn/status/2080067939152384306

Peter Yang har åbnet en lille skill mod 20 genkendelige AI-skrivevaner: falske kontraster, oppustet betydning, dramatiske fragmenter og resten af LinkedIn-maskinrummet. Den erstatter ikke en redaktør, men den kan i det mindste konfiskere modellens megafon.

Peter Yang@petergyang

I’m sick of reading AI slop, so today I’m open-sourcing my /no-ai-slop skill that removes 20+ slop patterns from any piece of writing. 📌 Get the free skill here: https://github.com/petergyang/no-ai-slop If you find it useful, please consider starring the repo so more people can find it. Why I built the skill: I use AI to edit my writing because it helps me fix spelling, grammar, and clarity. But even the best models keep producing the same slop that this skill removes: → Binary contrasts: “It’s not X. It’s Y.” → Throat-clearing openers: “Here’s what nobody tells you.” → Fake-profound endings: “The future isn’t coming. It’s already here.” Use this skill responsibly. I always write a first draft manually before iterating with AI on edits and I make sure to do another manual pass at the end as well. That’s in contrast to using AI to automate pumping out slop end-to-end. 📌 Read my full post for more on how I try to use AI responsibly to edit without giving into the dark side: https://creatoreconomy.so/p/use-my-no-ai-slop-skill-to-remove-20-ai-slop-patterns

Videoforhåndsvisning af no-ai-slop-skillen
♥ 3.406↻ 196💬 167🔖 7.009
https://x.com/petergyang/status/2079943830024188105

Nyhedsbonus

OpenAI lancerede Presence, en begrænset enterprise-tjeneste til voice- og chatagenter med politikker, guardrails, simuleringer, evals, godkendte handlinger og Codex-drevne forbedringsforslag. OpenAIs egen supportagent løser ifølge selskabet 75 procent af henvendelserne uden mennesker. Det er F2-territorium med en stor amerikansk motor under kølerhjelmen.

Alphabet leverede de mere jordnære AI-tal: Google Cloud voksede 82 procent, Gemini-API'erne behandler cirka 22 milliarder tokens i minuttet, og Gemini Enterprise bruges af næsten 90 procent af Fortune 100. De er stadig kapacitetsbegrænsede. Hype har fået en resultatopgørelse.