Now out to all Plus and Business users. Happy building!
Dagens Vibes — 5. september 2026
Astra-natten gav både raketbrændstof og reality checks: lange agentløb kræver stadig en manager, men spil, 3D og kodeportering flytter sig absurd hurtigt. Døgnets større signal kom fra matematikken, hvor Claude gjorde Fermats sidste sætning computer-tjekbar på 11 dage.
Fra X-feedet
GPT-6 Astra nåede Plus- og Business-brugerne i nat. Den officielle historie er 1,9× hurtigere Codex-arbejde, computerbrug og søgbar hukommelse på tværs af context windows; feedets egentlige historie er, hvad folk allerede fik den til at bygge.
Den vigtigste hands-on-læring er ikke et benchmark, men Matt Shumers “Manager Loop”: del arbejdet i faser, lad en manager-agent styre implementeringsagenten, og mål synlig fremdrift. Astra kan arbejde længe; den kan også fortabe sig i pynt. Selv AGI har brug for en mellemleder. Tragisk, men nyttigt.
A lot of people are asking how I pulled off these super long-horizon builds with Astra. Astra is extremely powerful, but by default it struggled with a task this difficult. I tested a bunch of approaches to get past this, and the one I landed on is something I'm calling the Manager Loop. It's basically a couple of tricks we used to use with much less capable models a couple of years ago, with a few new ideas layered on top. Turns out that when you put those together and apply them to Astra, its ability to do extremely difficult long-horizon tasks goes up dramatically. Here's how it works: 1. Launch an agent (I'm calling this one the "manager"). Chat with it about what you want to get done, and have it build a massive checklist of to-dos, then break that checklist into phases. 2. The manager then spawns a second Codex agent in a separate thread (the "implementer"). The two agents can message each other. 3. Put the manager in /goal mode, and tell it to run each phase on the implementer in /goal mode. 4. The manager messages the implementer: "/goal Complete phase one completely, extremely well." The implementer doesn't stop until that phase is done, then messages the manager back. The manager tells it to start phase two. They repeat until every phase is finished, completely autonomously. Why I think this works: over a long-horizon task, Astra tends to asymptote. It gets way further than previous models, but at a certain point it kind of just stops improving against the goal as quickly as it did before. It gets stuck in the minutiae, focusing way too much on small details, and overall progress stalls. The Manager Loop forces it to work piecemeal, one phase at a time. It's essentially how a human would steer a model, except the model is doing the steering for me. That's actually how this started. I was having the model write the checklist and break it into phases, and then I was doing the manager's job by hand. At some point I thought, "Wait, why can't I just get a separate AI to do this?" That's what unlocked full autonomy, which is super useful. A wording detail that seemed to matter: I ask for each phase to be done "extremely well," not "perfectly." Maybe I'm reading too much into it, but asking for "perfect" sent the model right back into the minutiae. "Extremely well" implies it's allowed to move on once it's good enough, and that worked better in my testing. One more trick that I think helps (this one is more of a hunch, but it was useful for me): have the implementer build a simple HTML page with the full checklist on it. The implementer checks boxes off as it goes and updates a counter, and the page has a chart of # of boxes ticked over time. Obviously the boxes aren't all equal, but it forces the model to notice things like "I haven't made progress in a while, time to move on". You can even put this in the prompt directly, like: "if you haven't ticked a box in X amount of time, move on". That helps a lot. I also ran 96 sub-agents at a time. You can change this in your Codex config (or just ask Codex to change it). This got me far better long-horizon performance than anything else I tried. I'll be sharing more in the coming days!
De stærkeste visuelle tests lukkede løkken mellem kode og billedgenerering: Anshu fik et 3D-spil på 45 minutter, mens Ethan Mollick gjorde Zork til et spilbart Three.js-actionspil med plot og puzzles intakte. “Lav en prototype” er ved at miste sin betydning, fordi prototypen allerede har boss fights.
dude GPT-6 Astra is some kind of turbo-AGI machine god for 3D games. It one-shot this in 45 minutes for hardly a couple % of my quota. I figured out how to get great graphics out of it. The trick is image gen. I'll share the process below.
This impressed me: I asked GPT-6 Astra to turn Zork (the classic 1977 text adventure) into a full 3D action-adventure game It kept the original plot & puzzles, added fight scenes, and built all of the characters & environments directly in Threejs Play: https://zork-underground-empire.netlify.app/
Kodearbejdet ser mere jordnært — og mere brugbart — ud. Paul Solt porterede en C++-raytracer fra Windows XP til Swift/Metal på en time. Men et offentligt KiCad-forsøg brugte 2 timer og 20 minutter samt 15 % af ugens 20×-kvote og endte med fabrikationsstop. Astra kan flytte software mellem årtier; kobberbaner er stadig chefen.
GPT-6 Astra ported my 3D RAY TRACER from Windows XP to iPhone in an hour. C++ → Swift → Metal. For fast prototyping, start with Light thinking.
Astra is *not* helpful for routing complicated PCBs 😐 I ran the BIGGEST public experiment testing its capabilities and here are the results, 2h20 on "/fast" and 15% of my 20x weekly limits later Details and files: https://github.com/jlcjak/astra_piNas Small 🧵with my thoughts
Døgnets tungeste resultat kom fra Anthropic: dusinvis af Claude-agenter formaliserede Fermats sidste sætning i Lean på 11 dage. Det er ikke en ny matematisk løsning, men 13 millioner linjer computer-verificerbart bevis og 29.500 nødvendige mellemteoremer. Den vigtige arbejdsform var en delt graf over delmål — Manager Loop, bare med 350 års teknisk gæld.
Checking that a major mathematical proof is correct can take years. Formalization—converting the mathematical reasoning into a form computer proof assistants like Lean can verify—can help. Last month, Claude completed the first formalized proof of Fermat’s Last Theorem, one of the most famous theorems of all time. This was a project experts thought would take many years. It is the largest Lean proof ever written. Fermat’s Last Theorem was first proven in 1995 by Sir Andrew Wiles, more than 350 years after it was conjectured. Our proof, which totals over 13 million lines of code, provides machine verification. More importantly, it proves over 29,000 other theorems that the proof requires, across many areas of math which had never before been formalized. We see this as a major step in the long process of firming up the core of mathematical knowledge, building on work from three centuries of mathematicians and hundreds of contributors to Lean and Mathlib. We are optimistic that AI-assisted verification of mathematical proofs will help reduce the burden of refereeing mathematics in an era where more proofs are being produced than ever before. You can read about the process on our Science Blog: https://www.anthropic.com/research/formalizing-fermats-last-theorem And see the complete proof on GitHub: https://github.com/anthropics/fermats-last-theorem
Nyhedsbonus
AI-infrastrukturens tal er stadig skrevet med en mistænkeligt stor tusch: britiske Nscale søger 3,5 mia. dollar før en mulig børsnotering — 1,5 mia. i konvertible lån og 2 mia. fra NVIDIA — kort efter en compute-aftale med Anthropic på cirka 45 mia. dollar. Agenternes nye superkræfter har en ret fysisk elregning.