GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday. We’re expanding preview access globally now.
Dagens hovedvibe: modellerne former sig til arbejdsroller. Sol som hurtig coworker, Fable som dyr specialist, og agent-workflows som den egentlige historie bag al model-fyrværkeriet.
GPT-5.6 Sol/Terra/Luna bliver dagens store model-anker. Feedet læser det som: Sol er ikke bare “smartere”, men et nyt arbejdsdyr til længere agent-loops.
GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday. We’re expanding preview access globally now.
Mitchells Sol-vs-Fable-billede er dagens mest præcise modelpersonlighedstest: hurtig, socialt brugbar coworker mod genial eneboer.
I had early access to 5.6/Sol for ~month. Sol is my default. It is faster, plans/judges just as good as Fable, and I think produces better overall work. I’ll reach for Fable still for highly targeted debug or performance work with clear reward functions. A cheeky way I describe Sol vs Fable to my friends is that Sol is a charismatic, efficient, talented coworker you’re jealous of. Fable is a genius recluse that is brilliant at its fixations but doesn’t go out, doesn’t date, and you don’t want to hang out with them much lol. Fable is undefeated at highly targeted debug/security/performance goals. It’s a sight to behold and I was never able to get Sol to push as hard in this category. I’ll keep using it for this. Sol is better or comparable at everything else, in my experience. Give it a shot, it’s hard to describe but it’s just more enjoyable to work with. (Disclaimer I have no financial ties to either lab, wasn’t paid for any of this.)
Dan Shipper sætter ord på skiftet: ikke “AI hjælper med tasken”, men loops der kører knowledge work mens mennesket passer systemet. Kontorarbejde, men med flere akvarier.
GPT-5.6 is the first model I’ve used that can reliably run whole loops of knowledge work, not just help with individual tasks. Your job shifts from doing the work to tending the system that does it. I use it to run loops that autonomously 1) do my email 2) help me find new hires for @every 3) keep me up to date on key decisions in all internal meetings / slacks @every 4) scan facebook marketplace to help me find furniture for my apartment
Cursor/xAI smider Grok 4.5 ind i værktøjskassen — endnu et tegn på at editoren er blevet model-markedsplads, ikke bare IDE.
We've partnered with SpaceXAI to train Grok 4.5. It’s our most powerful model yet and the first we've built for more than software engineering.
Eval-krisen fortsætter: OpenAI trækker anbefalingen af SWE-Bench Pro tilbage efter at have fundet cirka 30% broken tasks.
We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a leading coding eval.
TypeScript 7 er officielt ude: native Go-port, 8-12x hurtigere builds og mærkbart hurtigere editor-loop. Det er usexet på den gode måde.
Huge milestone for our team today: TypeScript 7 is now generally available--a native port that runs 10x faster. @typescript
Bun-in-Rust-posten er dagens store agentic engineering-case: en enorm rewrite drevet af workflows, test-suite og hærførte Claudes.
Rewriting Bun in Rust
Cloudflare Drop er den lille praktiske sidequest: smid en mappe i browseren og få et site live. FTP har taget solbriller på og lader som om det er nyt.
What do you mean by drop your folder? Did we just bring back FTP??? Gotta try this out, taking it for a spin tomorrow 👀
Introducing Cloudflare Drop Drop your folder in the browser and deploy it instantly on Cloudflare. Your website... milliseconds away from users on region: earth No account needed. Deployment is active for 60 minutes, then expires unless you claim it.
OpenAI foldede også benchmark-problemet ud i langform: agentiske coding-evals er svære at stole på, når opgaverne er høstet fra virkelige PR’er med skæve tests.
OpenAI finder 27–34% problematiske SWE-Bench Pro-opgaver og trækker anbefalingen tilbage. Benchmark-industrien får endnu en våd avis i ansigtet.
OpenAIhttps://openai.com/index/separating-signal-from-noise-coding-evaluations/Simon Willison peger på Bun-rewriten som case study i parallelle agent-workflows: conformance suite, adversarial review og procesfixes i stedet for bare kodefixes.
Kort, klar læsning af hvorfor Bun-historien er mere interessant som agentic engineering end som sprogkrig.
Simon Willison’s Webloghttps://simonwillison.net/2026/Jul/8/rewriting-bun-in-rust/