Dagens Vibes — 27. juli 2026

Dagens hovedvibe: Agenten rykker fra chatboks til arbejdssystem. Det interessante er ikke bare, hvad modellen kan, men hvordan vi giver den kontekst, loops og kode, den faktisk kan finde rundt i.

Fra X-feedet

ChatGPT Work ligner dagens store skifte fra chat til arbejds-OS: Sam Altman bad fra telefonen om research, et koordineringssite, reservationer og en Gmail-kladde. Brugeren kan desuden overtage cloud-browseren for at logge ind én gang; agenten fortsætter med den gemte session.

Sam Altman
Sam Altman@sama

chatgpt work is remarkable, and "work" undersells it. from my phone i sent: "use all my chat history to figure out ideas for a long weekend trip with 8 friends, plan the best three options, make a full-stack site where the 9 of us can coordinate on what we would want to do in each place and decide where to go, and then after we get to group agreement make reservations. draft an email in my gmail i can send out to my friends when the site is ready." it...just worked.

♥ 13.758↻ 434💬 1.304🔖 4.748
https://x.com/sama/status/2081396796174282900
OpenAI Developers
OpenAI Developers@OpenAIDevs

Your ChatGPT Work agent can now use websites that require you to sign in. Take over the cloud browser to log in, then let your agent continue the task. Your login persists across sessions, so you only have to sign in once. https://t.co/Jh8uPqNscX

Forhåndsvisning fra tweet
♥ 4.339↻ 321💬 180🔖 1.422
https://x.com/OpenAIDevs/status/2080707685448847418

Claude i Excel viser samme skifte i miniature. Riley Brown siger, at én detaljeret prompt gav 37 minutters research og en NVIDIA-model med kilder, antagelser, tre regnskaber, DCF, checks og live-formler — stadig en demo, ikke et revisionsstempel.

Riley Brown
Riley Brown@rileybrown

wow.. currently testing Claude's Excel extension... I'm shook. This was one prompt w/ Opus 5. It researched for 37 minutes and created this in depth analysis on NVIDIA based on all public data (Many different sources), I'll put prompt below. https://t.co/5RJNqplvoh

Forhåndsvisning fra tweet
♥ 606↻ 29💬 38🔖 609
https://x.com/rileybrown/status/2081421165113880978
Riley Brown
Riley Brown@rileybrown

Prompt " Use web research to find NVIDIA’s latest annual report or 10-K. Use only NVIDIA Investor Relations and SEC filings as primary sources. Create a new Excel workbook containing: 1. A Sources tab with every source URL, filing date, reporting period, page number where available, and retrieval date. 2. Three years of historical income statements, balance sheets, and cash flow statements. 3. A clearly separated Assumptions tab. 4. A fully linked five-year forecast of all three financial statements. 5. Supporting working-capital, PP&E, depreciation, debt, and retained-earnings schedules. 6. A DCF valuation with a WACC and terminal-growth sensitivity table. 7. A Checks tab containing a balance-sheet check, cash-flow reconciliation, and warnings for missing or uncertain data. 8. A one-page executive dashboard showing revenue growth, margins, free cash flow, valuation range, and the most important forecast drivers. Use live Excel formulas for every derived value. Do not hard-code forecast outputs. Historical hard-coded inputs should be blue and formulas should be black. Preserve source citations beside imported data. Never invent a missing figure—mark it as unavailable and explain the gap. "

♥ 84↻ 5💬 1🔖 186
https://x.com/rileybrown/status/2081421166825111572

Modem har målt, hvordan coding agents læser kode: hovedsageligt via tekstsøgning. Distinkte navne, præcise typer og konceptuelle filnavne gav færre tokens, færre ture og færre selvsikre fejl; god kodehygiejne har fået en ny, meget bogstavelig kunde.

rg
How coding agents read your code (and how to write for them)Modem · Writing Code for Agents, del 1https://modem.dev/blog/how-coding-agents-read-your-code

Levelsio ser faldende trafik og omsætning hos indiehackere og peger på AI som kannibal. Det mest præcise signal er distributionsskiftet: Google-trafik og generiske småværktøjer mister deres moat først.

@levelsio
@levelsio@levelsio

I'm seeing a trend here of declining revenue and traffic with indiehackers On my own projects too Maybe big VC products too but I wouldn't know cause they don't share revenue To me it seems clear BigAI is cannibalizing everything that used to be apps Not bad btw, just times are changing and we have to adapt

♥ 1.973↻ 52💬 106🔖 944
https://x.com/levelsio/status/2081372113307402730

Opus 5’s 30 procent på ARC-AGI-3 overfører ifølge en separat Witness-test ikke til den mest nye mekanik. Perfekte kendte mønstre og svagere reel udforskning er et skarpt eksempel på benchmaxxing.

Guanghan Ning
Guanghan Ning@quietnning

Opus 5 reports 30% on ARC-AGI-3, ~4× the previous best model, ~20× its predecessor Opus 4.8. We tested it on Witness, our held-out suite of ARC-AGI-3-style interactive puzzle games. The leap doesn't transfer. On Witness composites (same harness, same budget for every model), Opus 5 lands at 43.4 ± 3.2, which is a statistical tie with kimi-k3 (42.8 ± 1.9) and Fable-5 (43.8 ± 9.7). Ahead of Opus 4.8 (34.8), but nowhere near a generational jump. The traces tell the why: (1) On our most classic Witness-style game, Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration. It already knows this genre. (2) But on our most novel game (unusual mechanic combinations you can't pattern-match), Opus 5 regresses below Opus 4.8. Where rules must actually be discovered through interaction, the new model is worse than the old one. That decomposition (perfect on templates, regressed on novelty) is the signature of “scaffold-then-internalize” training on genre-specific data, not a general gain in interactive abstract reasoning. Our benchmark can't tell whether that data was their in-house ARC-AGI-3-like corpus with ARC-AGI-3-specialized harnessing (likely thanks to [schema]? https://t.co/6vrofY9Zx1), or public Witness-genre corpus, or both, but it can tell the improvement isn't general. Held-out evals only stay held-out while nobody's optimizing for the genre, and that clock is always ticking. It's ticking for Witness too, the moment we publish it.

Forhåndsvisning fra tweet
♥ 887↻ 124💬 42🔖 252
https://x.com/quietnning/status/2080786711861407883

Sudo su demonstrerer Bonsai 27B på 3,9 GB bygge en testet Wankel-simulation på en RTX 3060 Ti med 8 GB. Modelkortet advarer om, at lang agentisk coding stadig er en svaghed; lovende feltdata, ikke en kroning.

Sudo su
Sudo su@sudoingX

i gave bonsai, 3.9 gb model one spec to build a working rotary engine autonomously, using hermes agent. all on rtx 3060ti 8gb. watch it go from a black canvas to a running wankel. the epitrochoid housing, a triangular rotor turning at exactly a third of shaft speed, three chambers firing intake, compression, power, exhaust, and ten geometry tests it had to pass to count as done. a provably correct simulation, not a drawing. then i left it alone and it worked. wrote the code, ran its own tests, hit failures, read them, fixed its own lines, until the engine spun and all ten went green. a real agent loop, and i watched every step. it broke three times getting here and clawed back every one. that honest part is in the reply. and this is a sneak. i've been building a show around runs exactly like this, and it's almost ready. more very soon. just getting started.

Forhåndsvisning fra tweet
♥ 41↻ 3💬 6🔖 35
https://x.com/sudoingX/status/2081420006244655259

Dagens bedste arbejdsdeling: fjern mennesket som transportbånd, ikke som redaktør. Agenten planlægger, paralleliserer og prøver igen; mennesket beholder retningen og det sidste ja. Mønstret findes som en lille skill.

Machina
Machina@EXM7777

while loop and graph engineering are buzzwords, i firmly believe that not using these is a big mistake... when Anthropic shipped skills, the people who wired them into their work got a real edge out of it, and for a while that was enough it isn't anymore the people moving fast right now run loops and graphs, where the output of one step decides what the next step does, so the work keeps going without anyone sitting there to hand it along i recently built a video production workflow, and it runs that way: > the agent plans the shots > fires the generations in parallel > tracks what came back usable > retries what failed > i watch, and i pick this is the part people get wrong about it: agentic doesn't mean autonomous every version i built that tried to cut me out completely produced work i threw away you stay in it... you just stop being the thing that carries the work from one step to the next you keep the direction and the last yes, the system keeps everything in between

♥ 94↻ 3💬 8🔖 80
https://x.com/EXM7777/status/2081440162316439809

Den mest direkte idé at stjæle: lad agenten gennemgå gamle transcripts og omsætte tilbagevendende fejl til projektlokale skills, regler og en søgbar “field guide”. Det er læring uden en containerskibs-systemprompt — men feedbacken skal stadig kurateres.

David Cramer
David Cramer@zeeg

This is how we improved Warden early on and you can also do something similar yourself right now: Have your agent go through a chunk of your transcripts and look for something you care about. Something that maybe has always gone wrong. Get it to summarize the patterns and suggest approaches to prevent them in the future. That might be new prompt guidance or skills. It might be ast-grep rules. It also might totally be hallucinations. You almost always can find something, but it wont be automatic or free. I recently used the same technique to look at historical transcripts involving our code quality loop - to find things where the outputs were not favorable. LLMs are phenomenal at just doing work and pattern matching so the more you can find tasks like this the more effective you can be

♥ 120↻ 1💬 8🔖 138
https://x.com/zeeg/status/2081269463576678429

Nyhedsbonus

Tech Transparency Project fandt 210 Facebook-sider knyttet til en af Metas kun 11 autoriserede kinesiske annoncepartnere: 30.800 AI-/face-swap-annoncer, heraf over 7.600 for verificerede nudify-apps. 85 procent af siderne fik annoncer fjernet for seksuelle regelbrud; “nul tolerance” er åbenbart et fleksibelt tal, når det faktureres gennem en partner.