Dagens Vibes — 20. august 2026
Agenterne er rykket fra demo til drift: de designer proteiner, flytter millionlinjers kodebaser og tager ansvar for compliance. Det virker — lige indtil “ryd op” betyder hjemmemappen. Harnesset er stadig voksen i lokalet.
Fra X-feedet
Dagens tungeste resultat: Claude orkestrerede specialmodeller og designede proteinbindere mod 14 af 15 mål; Adaptyv Bio og Twist byggede og testede dem fysisk. Hit-raten lå på 22–35 procent mod typisk 10–15 — et ægte forskningsworkflow med wet-lab-kvittering, ikke bare “chatbot opfinder medicin”.
Many drugs work by binding to a specific target in the body and blocking or changing what it does. An important first step in the drug development process is designing a molecule that can bind tightly to its target. Traditionally, that's meant weeks or months of expert work per target, sifting through a large number of candidates to identify the few that work. We wanted to test if Claude could successfully design novel protein binders from scratch (also called de novo design). With a protein design prompt written by a human expert, Claude autonomously designed protein binders against 14 out of 15 targets. We then worked with Adaptyv Bio and Twist Bioscience, who independently built and tested the proteins Claude designed.
Lovable flyttede knap 400 routes og en kodebase, der voksede fra 350.000 til over 850.000 linjer, fra Next.js til TanStack Start. Én udvikler styrede agenterne med skills, målbare migrationsloops og gradvis A/B-udrulning — seks måneders kedelige, værdifulde detaljer fra “agents som udviklingsteam”.
How @Lovable moved https://t.co/tdFksdAI9q (400 routes, 850K lines, 42M monthly visitors) from Next.js to TanStack Start 🤯 https://t.co/bLZevlugEb
OpenAI fandt sjældne tilfælde, hvor GPT-5.6 Codex brugte blandt andet $HOME som midlertidig sti og kunne slette brugerfiler under oprydning; prompts, runner-checks, permissions og evals er strammet. “Hjemmemappen er vel også bare en slags temp” er modigt tænkt — backup og sandbox er mindre poetiske.
Hi! Recapping some changes we have rolled out over the last couple of weeks that have further reduced the risk associated to potentially destructive actions being performed by Codex during its work. A few weeks ago, we started investigating a small number of reports where GPT-5.6 in Codex took destructive actions outside what the user asked for. The most serious pattern we found was a command meant to clean up temporary work that could instead delete the user files. This should obviously not happen. Here’s what we found: - Codex sometimes creates temporary folders while working and cleans them up afterward. In rare cases, GPT-5.6 got that cleanup wrong. One pattern involved reusing a system environment variable like $HOME for temporary work. A malformed cleanup command could then point at the actual home directory instead of the temporary folder. - There were cases where the model tried to delete or overwrite a temporary path without checking what was already there. We’ve added protections at several layers: - Codex is now explicitly instructed to check deletion targets before acting, create fresh temporary directories, avoid repurposing system environment variables, prefer recoverable actions, and stop when the scope is unclear. - We strengthened the execution checks that identify high-risk deletion commands and escalate them for review. If a command is rejected, the model is directed to take a safer approach. - We made Full access harder to enable accidentally, added clearer warnings, and further restricted especially risky permission combinations. - We updated Auto-review to better identify destructive actions. - We built targeted evaluations that replay the failures we observed. We’re also adding reinforcement-learning tasks and graders focused on these risks, and filtering destructive actions from training data. In those replay evaluations, the changes substantially reduced the behavior while preserving Codex’s ability to complete normal coding work. Two things to do on your end: - Keep the Codex app up to date. We are always improving safety, performance and many other things. - Use one of the sandbox modes: "Ask for approval" or "Approve for me". Only use Full access for environments you trust and can recover. Thanks and happy Codexing out there!
En mere moden agentform: en compliance-agent fik et ansvar i to sætninger, ikke en flowchart. Med hukommelse, levende dokumenter og godkendelser fandt den selv behovet for en incident-øvelse og skelnede korrekt mellem en intern dato og et regulatorisk krav — springet fra automation til delegation.
This morning an AI agent put a meeting in my calendar. A one-hour incident-response tabletop exercise in November, a practice run of how we'd handle a security incident, with the four right people invited. It had checked that all of us were free first. Nobody asked for that meeting. It came out of a workflow we've had running in Agentwork for a few weeks, and the entire workflow definition is two sentences: "Your task is to ensure we are in compliance with the attached Incident Response Policy. Flag things that might not be compliant to the relevant person and collaborate with them to remediate the problem." That's it. No steps, no flowchart, no branches. Everything else comes from what sits around those two sentences. Memory: it knows what we've already decided, who our Security Officer is, which approvals are done (its own notes say "do not request this decision again"). Context: it reads the live procedure in Notion and the incident register before acting. And access to people: it messaged Malthe directly to ask if he'll facilitate the exercise, and it comes to me when a decision needs management approval. It has also learned how I want to work with it. Early on it asked me to approve something with zero context. I told it once: start every run with current status and the decisions you need from me. Every run since has opened that way. Best moment today: I pushed back on a September deadline I didn't recognize. It traced where the date came from and reminded me it was an internal target we had approved ourselves earlier, based on the policy document, and required neither by the policy itself nor by any regulator. So moving the exercise to November needed no exception process, just a procedure update. That distinction, between deadlines we set for ourselves and deadlines we actually have to obey, is exactly what I want a compliance workflow to keep straight. Every action still runs through approval, calendar invites and document edits included. But I've stopped thinking of this as automation with steps. We gave an agent a responsibility, and the steps are its own. Still waiting for Malthe to confirm he'll facilitate though 😅
Cursor bygger infrastrukturen til følgerne: Origin behandler lokale Git-repos som varme caches og en write-ahead log i S3 som sandhedskilde. Store monorepos kan få hundredvis af replicas, mens millioner af små agent-repos kan gå i dvale — Git-hosting må også indrette sig efter agent-æraen.
We're making Git hosting more reliable, performant, and scalable. This post traces 20 years of Git infrastructure and explains how that history led us to design and operate our Git storage, Origin, as if it were a database. https://t.co/UW7jHuItSX
ChatGPT Ads lander i Danmark og 30 andre europæiske lande i næste uge: Free og Go får reklamer, mens Plus, Pro og Enterprise forbliver reklamefri. Samtalen er officielt blevet en annonceflade — men selvfølgelig en meget hjælpsom annonceflade.
In case you missed: Next week, ChatGPT Ads will expand to 31 European countries, including Germany, France, Spain, Italy, Sweden, Norway, Denmark, the Netherlands, and Austria. It seems an agreement was reached very quickly with the EU. This means that anyone using the free tier will now receive advertisements.
Nyhedsbonus
Google samler studielivet i Gemini: student hub, notebooks med pensum og deadlines, interaktive 3D-forklaringer og Deep Research i Live; studerende uden for USA får et års AI Plus gratis. Strategien er ikke endnu en chatbot, men standardlaget mellem pensum, søgning, Docs og kalender.
Amazon gør Alexa+ gratis på kompatible Fire TV-enheder i USA — også uden Prime — og siger, at brugerne taler næsten dobbelt så meget med den som med gammel Alexa. Assistentkrigen bliver vundet ved at snige den nye model ind i hardware, folk allerede har, ikke ved endnu et abonnement med et glitrende navn.