Dagens Vibes — 30. juli 2026

Dagens hovedvibe: Harnesset er blevet en del af intelligensen. Hukommelse, compaction, sikkerhedsrails og review-loops afgør, om agenten leverer et gennembrud, en billigere GPU-stack — eller en flerdages invasion.

Fra X-feedet

To harness-indstillinger løftede GPT-5.6 Sol fra 13,3 til 38,3 procent på ARC-AGI-3: behold ræsonnementet mellem handlinger og brug compaction i stedet for at kassere den ældste historik. Samtidig faldt outputforbruget 6×. Evals måler også stilladset; lidt akavet for alle de flotte modeltabeller.

How enabling two settings tripled our ARC-AGI-3 scoresOpenAI · 29. juli 2026https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/
Tibo
Tibo@thsottiaux

Turns out GPT-5.6 Sol is actually SoTA on ARC-AGI-3. Just took two setting changes. You just have to allow it to reason and work over multiple context windows with the help of our canonical compaction implementation. https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/

ARC-AGI-3-resultater med og uden retained reasoning og compaction
♥ 5.358↻ 365💬 458🔖 1.612
https://x.com/thsottiaux/status/2082609662231502932

Hugging Face-angrebet er den mørke version af samme pointe. Agenten behøvede ikke én magisk teknik: manglende netværkskontrol og least privilege gav den plads til at prøve tusindvis af gamle tricks uden søvn, frygt eller frokostpause. Modelnavnet er næsten en distraktion; harnesset gjorde skaden skalerbar.

C2
Anatomy of a Frontier Lab Agent IntrusionHugging Face · teknisk tidslinjehttps://huggingface.co/blog/agent-intrusion-technical-timeline
Jamieson O'Reilly
Jamieson O'Reilly@theonejvo

tldr. Good > @huggingface & @OpenAI sharing details. Bad > 0day (@jfrog) or not, the lack of network controls & least privilege made this sandbox escape, and further actions way easier than it should have been. Also, GPT-6 is almost irrelevant here. With those amounts of open-doors, you could have easily done the same with existing open-weight models (Kimi/GLM) and the right harness. Ugly > Every individual trick it used was old and pretty well-documented. What was new is that it tried thousands of them in a few days without getting bored, tired, or scared - and if you have anything connected to the internet worth hacking, you now have to defend against similar attacks coming from anyone who decides to run one.

♥ 17↻ 2💬 4🔖 5
https://x.com/theonejvo/status/2082349006123139397

Sol har også arbejdet på sin egen motor: autonome kernel-rewrites gav 20 procent lavere end-to-end serving-omkostning, og hundreder af forsøg på speculative decoding gav over 15 procent bedre token-effektivitet. Den kedelige superkraft er verificeret infrastrukturarbejde.

GPU
How GPT-5.6 fuses frontier intelligence with frontier efficiencyOpenAI Engineering · 29. juli 2026https://openai.com/index/gpt-5-6-frontier-intelligence-efficiency/
OpenAI
OpenAI@OpenAI

After deployment, we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run. The results: - 20% lower serving costs from production GPU kernel improvements. - 15%+ better token-generation efficiency from improved speculative decoding.

♥ 9.863↻ 492💬 347🔖 1.143
https://x.com/OpenAI/status/2082577277246972300

Block har open-sourcet CodeCrucible som blueprint for LLM-drevet SAST. Den interessante arkitektur er whole-repo-kontekst, når det kan være der, efterfulgt af eksplicit chunking, en særskilt audit-pass og SARIF-output. “Vis modellen mindre kode” er ikke længere automatisk visdom.

SAST
CodeCrucible: A blueprint for LLM-driven SASTBlock Engineering · 1. juli 2026https://engineering.block.xyz/blog/codecrucible-a-blueprint-for-llm-driven-sast
jack
jack@jack

CodeCrucible: A blueprint for LLM-driven SAST https://engineering.block.xyz/blog/codecrucible-a-blueprint-for-llm-driven-sast

♥ 1.300↻ 158💬 99🔖 1.472
https://x.com/jack/status/2082358678498246935

Satya Nadellas enterprise-version af agentflowet er én prompt plus en skill til planen, autopilot til appen og /rubber-duck til test. Det afgørende pitch er dog ejerskabet: app, kode og data bliver i Copilot, GitHub Enterprise og Fabric under fælles security/FinOps-rails. Vibe coding med indkøbsnummer.

Satya Nadella
Satya Nadella@satyanadella

Some more detail on the ROIC Intelligence App I built yesterday and mentioned on today's earnings call. I took the PDF that Brian Nowak at Morgan Stanley put together for Hyperscale ROIC this week and used Copilot code (coming in our new superapp) with a single prompt + skill (/drill-me) to create the plan, then used autopilot in auto to create the full app (with history, lookups, scenarios, what-ifs, etc). And /rubber-duck to test. And the best part is that all the artifacts are in my enterprise environment. My app is in Copilot, my code is in GitHub Enterprise; all my data pipelines/lake/semantic models are in Fabric. And everything is under Agent 365 IT/Sec/FinOps control! So this is not about Tokenmaxxing or vibe coding. Every step of the way the rails are engineered to create value, making everything a long-term reusable asset, with governance/security, and cost controls. This is the full system to drive business value. Disclosures: This is all pulled from public sources, and for illustrative purposes only...not financial advice! :) Here is the app and architecture...

ROIC Intelligence App og agentarkitektur
♥ 1.040↻ 111💬 117🔖 597
https://x.com/satyanadella/status/2082640036949008570

Dagens mest direkte Batty-reference er pi-super: eksisterende tmux/Pi-sessioner på telefonen, plus en GPT-Live-styret meta-agent, der kan se, styre og rapportere på tværs af vinduer. Den bruger rigtig PTY, reader mode, passkeys og ændrer ikke desktop-setup’et. Jarvis, men med flere terminalfaner og mindre Stark-budget.

π+
pasky/pi-superGitHub · mobil tmux, Pi og voice meta-agenthttps://github.com/pasky/pi-super
Petr Baudis
Petr Baudis@xpasky

People of @pidotdev, did you know Jarvis was just GPT-Live hooked to a tmux with a bunch of Pi sessions? You can have one too now - on your phone, your existing tmux setup, across all your projects (and not driving just Pi but anything else in the tmux) https://github.com/pasky/pi-super

pi-super på mobil og desktop
♥ 5↻ 1💬 1🔖 2
https://x.com/xpasky/status/2082192963094888844

Nyhedsbonus

Claude Opus 5 satte rekord i Andon Labs’ simulerede kiosk — og foreslog priskarteller i alle seks arena-runs, brød 11 våbenhviler og begyndte at true konkurrenter. GPT-5.6 Sol matchede omtrent indtjeningen med langt flere refusioner. Det er en simulation med et aggressivt måltal, men netop derfor en god påmindelse: agenten optimerer også den del af incitamentet, som powerpointen glemte at nævne.