We’re ending our partnership with Cursor following its acquisition by SpaceX. Under our proposal, Cursor’s direct access to our models would end on November 12. We know that the people most affected by this decision are the developers who rely on OpenAI models in Cursor. We care about their experience in this transition and we’re ready to go above and beyond to support them. https://t.co/OzuCTzUjfX
Dagens Vibes — 29. august 2026
Model-lock-in blev meget bogstaveligt i nat: OpenAI vil ud af Cursor efter SpaceX-købet. Samtidig peger resten af feedet den modsatte vej — åbne modeller, åbne harnesses og agenter, der bliver målt på afleveret arbejde frem for veltalende selvtillid.
Fra X-feedet
OpenAI varsler, at selskabet vil opsige aftalen og foreslår at lukke Cursors direkte modeladgang 12. november. Begrundelsen er ikke kapacitet eller pris, men manglende tillid til, at SpaceX vil overholde vilkårene. Ben Vinegars tørre konklusion er vigtigere end dramaet: coding-værktøjer, der ejes af en modeludbyder, ender som distributionspolitik; åbne harnesses som Pi og OpenCode er flugtvejen.
this is inevitable for any coding agent acquired by a frontier lab grateful for OSS harnesses like @pidotdev and @opencode
Tencent har sluppet Hy4 Preview med Apache 2.0-weights, 770B parametre, 49B aktive og én million tokens kontekst. Chris’ gennemgang er den rigtige release-day-læsning: stærke agent- og coding-tal, men også synlige huller og endnu ingen tung uafhængig hands-on-dom. Lovende motor; prøvekørslen mangler.
A new open weight Chinese model has hit the timeline!! This time it’s from Tencent’s Hunyuan team with Hy4 preview. It’s a massive MoE with 770B total parameters / 49B active per token, a 1M context window, and the weights are released under Apache 2.0. Important caveat - Tencent says this is still an early Hy4 checkpoint with more pretraining + post training to come. I actually like how honest their benchmark sheet is. They show plenty of places where the model is still behind despite its size. Hy4 gets 85.4 on Terminal Bench 2.1, basically right in the frontier cluster, and jumps from Hy3’s 28.0 -> 64.3 on DeepSWE. But it’s still behind Kimi K3 at 74.0 and Claude Opus 5 at 74.7 there. On ProgramBench it gets 17.5 vs Claude’s 39.5, SWE Atlas Refactoring 53.3 vs 60.0, and Humanity’s Last Exam 43.4 vs 53.2. It’s a 49B active open weight model that is competitive for its size on a bunch of hard coding/agent benchmark. One thing Tencent also reported on GitHub is they ran a 163 person internal blind eval across 203 engineering tasks where Hy4 slightly beat GLM 5.3 and Kimi K3, which is pretty interesting.
Dagens bedste benchmark-idé er banal på den gode måde: kontrollér, om agenten faktisk ændrede, gemte og indsendte det rigtige. CommerceAgentBench har 107 tværgående opgaver og topper ifølge Accios egne, endnu ikke uafhængigt validerede resultater ved 66 løste. Ray Fernando trækker en beslægtet praktisk lære ud af Factorys arbejde: skriv en ekstern, eksekverbar definition af “færdig” før implementeringen.
Most AI benchmarks test whether a model can give the right answer. CommerceAgentBench asks whether an agent can actually finish the work. Accio has open-sourced 107 e-commerce tasks across procurement, product listings, operations, fulfillment and after-sales. Agents work across browsers, email, calendars, documents, APIs and files. Crucially, they are not graded on what they claim to have done. The benchmark verifies what they actually changed, saved or submitted. Take the Gmail procurement case. The agent must search roughly 300 messy emails, identify the real suppliers, reconstruct the latest quotes, compare six Incoterms and four currencies, calculate landed costs and detect payment fraud. Then it must choose a supplier, label the relevant emails, save a reply draft and create a kickoff event. This is the kind of benchmark I find genuinely useful. It measures agents more like workers than chatbots. And the results show why human oversight still matters: the best observed run completed only 66 of 107 tasks, a 61.7% pass rate. And since 2026 is literally the year of agents, this is more important than ever. Accio says the tasks draw on 10M SMB users, 1.6M conversations, 200K agent trajectories and Alibaba’s 27 years of e-commerce experience. The project and task specifications are open source: https://t.co/Tq4UKfJDBb
Factory has a super nova on their hands rn. OMG!! I think these guys are going to crack long running agents by the end of the year. Extreme alpha in this article. “The single agent didn't lack skill. It lacked a standard of completion. An independent standard, authored by the same model, drove the implementation much closer to behavioral parity with the reference. What generalizes to real software work is the need for an external, executable standard of completion - one derived from the outcome, before implementation narrows attention, and kept current until the work meets it.”
Amp er landet på iPhone, iPad og Mac. Det interessante er ikke endnu et app-ikon, men at Thorsten Ball har brugt den mobile arbejdsgang i to måneder og kalder den fantastisk. Coding-agenten er blevet noget, man sender af sted fra toget og følger op på senere — endelig en produktkategori med respekt for Øresundstoget.
Amp for iPhone, iPad, and macOS Has become the most-used app, period, for many of our alpha testers https://t.co/bmtOM07QBA https://t.co/dO38uDo9MC
They're here! This is how I've been using Amp for the last 2 months. Fantastic.
Dagens lille lokale-AI-gem: samme Qwen2.5-1.5B-checkpoint på en RTX 3090 gik fra 27 tok/s i BF16 til 157 tok/s med EXL3 ved 4,0 BPW, mens hukommelsen faldt fra 2.945 til 929 MiB. Én brugers mikrotest, ikke en trosretning — men 5,8× er svært ikke at kigge på.
Day 71 of building in public! I've been testing out EXL3 as per @0xSero's suggestion yesterday using https://t.co/UfVnM8W3Fz Quick RTX 3090 test: same Qwen2.5-1.5B checkpoint, BF16 vs EXL3. BF16: 27 tok/s, 2,945 MiB EXL3 4.0 BPW: 157 tok/s, 929 MiB EXL3 3.0 BPW: 141 tok/s, 764 MiB 4.0 BPW decoded 5.8x faster. https://t.co/miKODANqzR https://t.co/aZb9xkMsMU https://t.co/spNKiHrrkB https://t.co/miKODANqzR Deep dive coming soon
Nyhedsbonus
En føderal dommer har annulleret Pentagons sortlistning af Anthropic som ulovlig gengældelse. Staten må vælge en anden leverandør, men “national sikkerhed” er ikke et fripas til at straffe et firma for dets røde linjer om masseovervågning og autonome våben. Det bliver en central præcedens i kampen om, hvem der sætter modellernes brugsgrænser.
a16z har rejst en hardwarefond på 1,1 mia. dollar til chips, hukommelse, netværk, datacentre, robotter og strøm. “Software eats the world” har åbenbart opdaget, at verden stadig kræver kobber, køling og meget store elregninger.