Introducing Claude Opus 4.8: it builds on Opus 4.7 with sharper judgment, more honesty about its own progress, and the ability to work independently for longer than its predecessors. Available today at the same price.

Dagens feed var én stor Opus 4.8-støjsky, men signalet er workflowet rundt om: flere subagenter, hårdere review og bedre øjne på runtime. Koden får fabrikslys.
Opus 4.8 er ikke bare “ny model go brrr”. Det interessante er ærligere fremskridt, effort-knapper og Dynamic Workflows: Claude der planlægger, splitter ud i parallelle subagenter og prøver at få dem til at modbevise hinanden. Sund paranoia som produktfeature.
Introducing Claude Opus 4.8: it builds on Opus 4.7 with sharper judgment, more honesty about its own progress, and the ability to work independently for longer than its predecessors. Available today at the same price.

More on how the orchestration works, and what early users have built with it, on the blog: claude.com/blog/introduci…
Research preview: tens to hundreds of parallel subagents, resumable runs, adversarial checks og et Bun Zig→Rust-port med 99,8% tests på 11 dage. Ja, tokenmåleren griner ondt.
claude.comDer var også sund skepsis i feedet: benchmark-sammenligningerne er efterhånden så tæt på vibes, at man skal kigge på egen harness før man udnævner en vinder.
Anthropic did a big strategic error. Normally they compare their models with their old models. Instead today, now that everybody knows how strong GPT 5.5 is at coding, they put it in the mix, basically showing all their customers that the benchmarks can't be trusted. https://t.co/up73bHAfen

Det mest nyttige mønster i dag er ikke endnu en model-switch. Det er struktur: faste “pulse”-tråde, plan-docs som pseudo-kode, call stacks og parallelle review-agenter. Kedeligt? Ja. Virker? Også ja, den værste kombination.
codex power-user best practices: - a few different "pulse" threads that run every morning to check the status of stuff i care about, e.g. proof metrics, all company meetings, etc - a "log" thread for ongoing, everyday activity i want to track. for me that's my "writing log"—want to track what im writing about, and working on from piece to piece - "inbox" thread which gathers all of my emails right now, but eventually will pull in all of the main important things from each "pulse" thread - "router" thread which knows about all of the other threads, and is also hooked up to my email to push emails into each thread appropriately

my "plans" largely look like pseudo code composed of mostly types/interfaces, how they compose, and their boundaries ive recently started including call stacks - been very helpful for both me and agents when implementing https://t.co/SLrYX3ywqc

introducing thermos in cursor a deep security/correctness audit and a harsh code quality audit, run in parallel on your branch, synthesized into one prioritized list https://t.co/YKi2QeRGSA
Chrome DevTools for agents 1.0 er en af de mere praktiske nyheder: coding-agenter skal ikke bare skrive React-komponenter i blinde og håbe. De skal se browseren, netværket, konsollen og performance-sporet.
AI coding agents can write code, but they can't see if it actually works. Chrome DevTools for agents 1.0 fixes this. The stable release brings powerful browser debugging, emulation, and automated audits to your AI assistants via our Chrome DevTools MCP server. 👁️ Give your agent eyes on the runtime → goo.gle/42K7Rrl #GoogleIO

Stabil MCP-server til live Chrome-debugging, screenshots, console/network, performance traces, audits og browser automation for agenter. Endelig lidt syn i mørket.
github.com/ChromeDevToolsMidt i frontier-råberiet kom der også gode “små” nyheder: Liquid skubber tool-calling ned på laptop/telefon, Google gør Nano Banana 2/Pro bredt tilgængelig, og ByteDance/BAGEL-sporet lugter af færre specialmodeller per opgave.
Today, we're releasing LFM2.5-8B-A1B, a device-optimized model designed to power real-life applications on phones, laptops, PCs, robots, and fast & lightweight server-side use-cases. > 8B MoE, 1.5B active > Expanded 128K context > LFM2.5 flagship hybrid MoE architecture > Trained on 38T tokens + large-scale RL > fast, reliable tool calling, punching above its weight, comparable to models with up to 4x its size > customizable on a single GPU for any specialized task > LFM2 open-weight license 🧵

Nano Banana 2 and Nano Banana Pro are now generally available. 🍌 Our best image generation models are now available to developers and businesses in @GoogleAIStudio and via the Gemini Enterprise Agent Platform API. https://t.co/5NIDpxNrK4

ByteDance just open-sourced one of the most capable multimodal models out there. BAGEL does image generation, editing, style transfer, and visual understanding - all in a single 7B parameter model. Apache 2.0 licensed! One model. No switching between specialized tools. Amazing
Liquid AI’s 8B/1.5B-active edge model har 128K context, llama.cpp/MLX/vLLM/SGLang support og sigter direkte på lokale tool-calling-agenter.
Liquid AIPeter Steinberger byggede octopool efter endnu et møde med GitHubs rate limit: en Cloudflare Worker der pooler PATs/GitHub App-installationer bag en delt read-cache og drop-in gh-shim. Præcis den slags usexet infrastruktur der gør agent-flåder mindre hysteriske.
Hit GitHub's rate limit one too many times, so I built octopool: a Cloudflare Worker that pools your team's PATs + GitHub App installations behind a shared read cache. Self-host on Cloudflare. Drop-in gh shim. octopool.dev
Cloudflare-hostet read relay/cache til GitHub-adgang på tværs af teams og agenter. Ikke glamourøst. Meget brugbart.
octopool.devUden for feedet er det samme hovedtema: Anthropic prøver at sælge reliability og workflow, mens Google I/O pumper agent-flader ind i alt fra Search til Android-appbygning.
TechCrunchs korte læsning: 41 dage efter 4.7, samme pris, mere fokus på usikkerhed/ærlighed og en swarm-of-subagents feature som det egentlige produktgreb.
TechCrunchThe Verge opsummerer Gemini 3.5, Omni, Spark, vibe-codede Android-apps, Universal Cart og Pics. Google vil tydeligvis eje hele “AI som hverdagsoverflade”-laget.
The Verge