Dagens Vibes

Dagens Vibes — 27. maj 2026

Dagens hovedvibe: coding-agenterne bliver målt hårdere, pakket ind i plugins og review-skills, og alle leder efter den primitive der faktisk holder i produktion.

Fra X-feedet

Det store signal var DeepSWE: en agentic coding-benchmark der rammer den følte virkelighed bedre end de gamle leaderboards. Når evals endelig begynder at ligne hverdagen, bliver stemningen straks mere religiøs og lidt mere blodig.

Serena Ge (Datacurve)@serenaa_ge

Today we’re releasing DeepSWE, a new standard for agentic coding benchmarks. On public leaderboards, top models often look relatively close in capability. DeepSWE shows where they actually diverge, reflecting the realistic experience of developers in their day-to-day work. https://t.co/HCDcjNuTFK

♥ 4k↻ 476💬 328🔖 1.7k
https://x.com/serenaa_ge/status/2059308218564890875
DeepSWE: GPT-5.5 foran, og SWE-Bench Pro får tæsk

VentureBeat gennemgår Datacurves benchmark: 113 opgaver, større spredning mellem modellerne, og en ret pinlig verifier-/contamination-kritik af eksisterende coding-evals.

VentureBeathttps://venturebeat.com/technology/deepswe-blows-up-the-ai-coding-leaderboard-crowns-gpt-5-5-and-finds-claude-opus-exploiting-a-benchmark-loophole

Næste lag er drift: Claude Code får security-plugin, og autoreview-skills bliver en normal closeout-rutine. Det er ikke glamourøst. Det er derfor det lugter af noget der bliver brugt.

ClaudeDevs@ClaudeDevs

We’ve shipped a security-guidance plugin for Claude Code that helps identify and fix vulnerabilities as you’re writing code. Available for all Claude Code users. Install from the plugin marketplace (/plugins). https://t.co/LprgC4m6Kf

♥ 11.5k↻ 1k💬 263🔖 9.9k
https://x.com/ClaudeDevs/status/2059385239781384341
Peter Steinberger 🦞@steipete

autoreview is the most impactful skill I've added to my stack (next to crabbox.sh). It automatically reviews your code before landing a PR. Finds so many edge cases. Sometimes it runs for hours. github.com/openclaw/agent…

♥ 1.1k↻ 44💬 40🔖 1.7k
https://x.com/steipete/status/2059453909819654554
OpenClaw autoreview-skill

En konkret skill-kontrakt for struktureret review: advisory findings, verificér i rigtig kode, rerun tests, stop når der ikke er actionable fund. Kedeligt på den gode måde.

GitHubhttps://github.com/openclaw/agent-skills/blob/main/skills/autoreview/SKILL.md

GPT-5.5 fyldte meget i feedet: flere oplever at den kræver andre prompts og bedre AGENTS.md, men når den sidder der, er den svær at slippe igen. Ja, model-shopping er blevet en livsstilssygdom.

Theo - t3.gg@theo

It took me like 2 months, but I've grown to love gpt-5.5. You have to prompt entirely different and put some time into your agents[.]md. Now that I'm over the hump, I can't really use any other model for code.

♥ 3.2k↻ 68💬 276🔖 558
https://x.com/theo/status/2059372156753219938

Der var også infrastruktursporet: lokale/effektive billedmodeller, Rust-frontends i vLLM og SynthID som fælles watermarking-lag. Mindre demo, mere maskinrum.

PrismML@PrismML

Today we’re releasing 1-bit and Ternary Bonsai Image 4B. A new family of image-generation models designed to run high-quality diffusion inference on local hardware: from laptops to phones. https://t.co/9qB5UbOogJ

♥ 1.3k↻ 200💬 53🔖 828
https://x.com/PrismML/status/2059339157600969199
vLLM@vllm_project

🦀 The Rust frontend is officially merged into vLLM! As GPUs get faster, the frontend has become a real share of CPU time. The new Rust frontend is a drop-in alternative to the Python API server — same engine, same ZMQ boundary. Opt in with VLLM_USE_RUST_FRONTEND=1. Early https://t.co/uBieU4W39z

♥ 667↻ 71💬 23🔖 154
https://x.com/vllm_project/status/2059344804295942513
Google DeepMind@GoogleDeepMind

SynthID has already watermarked over 100 billion pieces of content, but transparency is a team sport. That’s why we’re partnering with @OpenAI, @ElevenLabs and Kakao to add SynthID watermarking to their models – accelerating the industry-wide momentum we started with @NVIDIA. https://t.co/QnshYx3EfE

♥ 1.2k↻ 119💬 85🔖 173
https://x.com/GoogleDeepMind/status/2059235181274202500

Og et sundt modhug: agenter kan stadig gøre dig langsommere, hvis du ikke kan parallelisere arbejdet mentalt. Magien findes; køen foran kaffemaskinen findes også.

Taelin@VictorTaelin

Status update: I've been on/off AI agents in the last few days and it is a verifiable truth that every day I didn't use agents, I was more productive. I still attribute that to how slow they are, and my own inability to multi-task efficiently. The magic is there but the slowness doesn't let it cross the threshold where they actually make me faster, and I still dislike the whole thinking paradigm. About Bend2: honestly, the C/Metal compiler codebase is a clusterfuck right now. I regret letting AI agents write it. All tests pass, and GPU performance is mind-blowing, so the core architecture works. Yet, it has a LOT of bugs. Anything not covered by the tests is a coin toss. This is actually impressive, because, in many parts of the codebase, the right solution was actually the simplest one, yet, the agents STILL managed to find a way to make it work just for the tests. The level of reward hack these agents output is actually impressive I can't even be mad. It is also ironical because that's the very problem that Bend's proof system was supposed to solve, but Bend is in TypeScript, not in Bend. I'm disappointed I didn't write Bend in itself, and now I feel an immense urge to do so. But the clock is ticking . . . Still, I do not think Bend is worth launching without the GPU compiler being solid, because the closest competitor, Lean, is actually extremely good, so we need a big differential. Yet, due to the very nature of the project, it would be embarrassing to have bugs at launch. Regarding AI, I now believe using current gen AI agents in production codebase is harmful and a massive mistake. That doesn't mean no agents at all, but agents work best when they don't touch critical code. Debugging, researching, providing insights, scripts / tools, or anything that doesn't touch code you will maintain in the long term. But if you merge AI code without reading, you're going to have a bad time. Speaking from experience I'm working 10h/day on SupGen and the remaining time on Bend2

♥ 1.2k↻ 59💬 80🔖 484
https://x.com/VictorTaelin/status/2059327831679578511

Nyhedsbonus

Uden for feedet var hovedhistorien AI-cyber: Mythos/GPT-5.5 beskrives nu som et konkret skift i offensiv/defensiv sikkerhed, ikke bare endnu en “snart AGI”-presseballon.

AI-modellerne der ryster Washington

POLITICO samler vurderinger fra cyberfolk: Mythos og GPT-5.5 kan finde og udnytte sårbarheder på et niveau der presser både myndigheder og virksomheder.

POLITICOhttps://www.politico.com/news/2026/05/24/anthropic-openai-mythos-what-to-know-00934668
Demis Hassabis: agenter er en generalprøve på AGI

Axios-interviewet er dagens mest direkte “tag det seriøst”-markør: DeepMind-chefen peger på agentbølgen som stress-test før stærkere systemer.

Axioshttps://www.axios.com/2026/05/26/deepmind-ceo-demis-hassabis