Dagens Vibes — 12. august 2026

Agenten flytter ud af chatvinduet og ind i maskinrummet: Grok Bot får sin egen computer, NVIDIA bygger et billigt udførelseslag, Claude designer selv sine workflowgrafer — og forskere opdager, at krypteret reasoning ikke var helt så krypteret, som brochuren lovede.

Fra X-feedet

Grok Bot er OpenClaw-idéen pakket som et massemarkedsprodukt: hver bot får sin egen cloud-computer, logger ind i dine værktøjer, lærer rutiner ved at kigge med og kan koordinere med andre bots. Det er agenten som kollega i stedet for chatvindue — inklusive den lille detalje, at man nu skal stole på Musk med sine loginoplysninger. Modigt.

Lenny Rachitsky
Lenny Rachitsky@lennysan

I got early access to Grok Bot and I'm hooked. I haven't been this excited about a new AI product in a while. It's like OpenClaw, but super easy, reliable, and less scary to use. I think this will be a huge new product line for Cursor/Grok/SpaceX. I've already found so many ways to use it that have meaningfully made my life better: 1. Matchmaking people looking for jobs with companies who are hiring (see below) 2. Auto-replying to support emails (saves me hours!) 3. Scanning my credit card statements and finding recurring subscriptions to cancel 4. Sending me (really good!) briefs for upcoming podcast guests See below for my actual set of agents that I've been using and chat with daily. Great work on this team Grok Bot. (I'm not an investor in this, nor do I have any ties to this product/company. I'm just a fan!)

Medie fra @lennysan
♥ 2947↻ 154💬 145🔖 1885
https://x.com/lennysan/status/2087241423792087518
BOT
Grok Bot — officiel introduktionEgen cloud-computer, lærte rutiner og teams af botshttps://x.ai/news/introducing-grok-bot

Dagens mest opsigtsvækkende paper er et reelt arkitekturhul: krypterede reasoning-blokke kan flyttes mellem brugere, sessioner og modeller, hvorefter en svagere model kan lokkes til at dekryptere dem. Forskerne fandt 367 persondatafund og 182 credentials i 315.320 offentligt delte blokke. Det er lossy dekryptering, men problemet er ikke akademisk pynt.

elie
elie@eliebakouch

insane paper, they extracted reasoning from frontier models by asking less safeguarded models (luna, haiku, ...) to decrypt and output the decoded version of the encrypted reasoning block. while reading this post or the paper, keep in mind this is a somewhat "lossy" decryption, but they have experiments to test it and it seems pretty accurate some stuff i found really interesting (much more in the paper and especially the appendix): - they were able to extract tokens and personal information from publicly available traces with this method - gpt models are trying to be very efficient in reasoning (was already in an apollo research report), extremely short sentences like "need maybe not need." "avoid maybe not" (they don't have 5.6 sol in this section tho!) - they computed perplexity on hidden reasoning with different open models and conditioned them on closed model reasoning as well > K3 seems to be more similar than other open models to sol and opus > result that makes me a bit sus is that they say they attribute higher average perplexity to their own CoT than the CoT of other models which i find quite surprising (except K3 and K2.7) but figure 26 showing median doesn't really show that > also surprising: all the open models score very high perplexity on gpt traces (see screenshot in thread) - models sometimes have chain of thought in other languages like chinese, russian or japanese. since it's lossy decoding this might not be super accurate but still interesting - models are mostly honest in their CoT on the task they tried and there are examples of them scheming and realizing they could cheat. also interesting that these things are not often transparent in the summary that users have access to - summary length doesn't really scale with the number of reasoning tokens really great work, maybe one thing is that i would have loved to see more details on cot of recent models like fable (i don't think paper have that) and 5.6 sol, lot of the observation here are based on older model cot, which is still valuable but don't automatically transfer to more recent model

Medie fra @eliebakouch
♥ 828↻ 64💬 26🔖 570
https://x.com/eliebakouch/status/2087179305474298162
PAPER
Stealing Reasoning Traces from Proprietary LLM APIsAngreb, målinger, datalæk og foreslåede mitigationshttps://arxiv.org/abs/2608.09867

NVIDIA går efter agenternes billige udførelseslag med Nemotron 3.5 Lightning: 30B MoE, kun 3B aktive parametre, åbne weights, data og recipes. Den interessante del er systemdesignet: frontiermodellen planlægger, Lightning tygger de tusind rutinekald, og Switchyard router mellem dem. Sol til dommen, lommelygte til skruerne.

Jensen Huang
Jensen Huang@JensenHuang

Lightning strikes for continuous and long-run agents! Nemotron 3.5 Lightning is smart, fast, efficient and open.

Medie fra @JensenHuang
♥ 7293↻ 695💬 451🔖 551
https://x.com/JensenHuang/status/2087184542050496763
NVIDIA
Nemotron 3.5 LightningÅben 30B MoE-model til agenternes udførelseslaghttps://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/

Anthropics Dynamic Workflows-cookbook vender orkestreringen på hovedet: Claude skriver selv workflow-scriptet og kan starte op til 1.000 subagenter. Agenten udfører ikke bare planen; den designer den midlertidige maskine, der udfører planen. Meget Batty-kompatibel galskab.

Daniel San
Daniel San@dani_avila7

Anthropic shared a Dynamic Workflows cookbook worth reviewing Bookmark this one Claude writes the orchestration script and spawns up to 1,000 subagents per run https://t.co/1xDUs5OMEk

♥ 746↻ 72💬 12🔖 1585
https://x.com/dani_avila7/status/2087026930323247306
CODE
Anthropic Dynamic WorkflowsCookbook: genereret orkestrering og op til 1.000 subagenterhttps://github.com/anthropics/claude-cookbooks/blob/main/claude_agent_sdk/08_Dynamic_workflows.ipynb

ChatGPT, ChatGPT Work og Codex lander som officiel Linux-desktopapp i preview til Ubuntu, Debian og Fedora på x64 og ARM64. Linux er dermed ikke længere det vindue, OpenAI glemte at montere i huset.

OpenAI
OpenAI@OpenAI

Now in preview: The ChatGPT desktop app for Linux. Use ChatGPT, ChatGPT Work, and Codex where you already work and build, with your projects and browser workflows on supported Linux systems. https://t.co/OtsPt5N5QC

Medie fra @OpenAI
♥ 9311↻ 740💬 609🔖 1029
https://x.com/OpenAI/status/2087231350134980830
LINUX
ChatGPT desktop til LinuxPreview, understøttede distributioner og Codex-integrationhttps://openai.com/codex/

EU’s AI Act gør detektion af AI-genereret indhold aktuel, men Deedys gennemgang punkterer den magiske detektor: tekstvandmærker er statistiske, først omkring 400 tokens når SynthID cirka 90 procent sande positiver ved under én procent falske. Proveniens er vigtigt; en usynlig moralsk stregkode i hver sætning findes stadig kun i PowerPoint-land.

Deedy
Deedy@deedydas

Deep dive into AI text-watermarking and what EU's AI Act actually mandates about AI detectability. How it works: The SynthID paper from Demis and team in Nature is probably the best primer in the technology. Generally, most approaches involve sampling next tokens statistically differently with some random key while not distorting the semantics of the output. Detecting text watermarks can never be perfect. Longer passages are always easier to detect, only at ~400 tokens do you get a true positive of rate of ~90% when your false positive rate is <1%. Because small changes in the text can change whether the algorithm thinks it's AI or not, other approaches include semantic data in their watermarking scheme too. Industry status on the EU regulation: EU's AI Act Article 50(2) requires AI generations from all modalities to be AI-detectable, with practical implementation details still being finalized. The main exception is "AI systems perform an assistive function for standard editing or do not substantially alter the input data" although how they would distinguish between these cases is unclear. Article 113 says "It shall apply from 2 August 2026." Claude declared that they will comply. Gemini already uses SynthID for their text outputs (even though no production public detector exists for text outputs). OpenAI had developer watermarking a while ago but reportedly shelved it when 30% of their users said they'd use the product less. However, their support page says "our goal is to expand provenance signals to all modalities including text" (!) I've always maintained I think it's critical to know whether text comes from AI or a human (and thus support tech like Pangram). Judging by the comments on X, I might be in the minority. This comment sort of embodies the backlash: "I'm paying it to write for me in a way people and machines wouldn't detect it is AI written" In other words, users believe if their ideas go into using AI for expressing them, *they* should be given credit for the idea and not penalized for using AI. I'd counterclaim that penalization and provenance are distinct, and provenance is extremely critical because a) without provenance, you are more likely to be judged poorly and falsely accused of using AI when you didn't b) just like with food, fundamentally readers deserve the right to know where the text came from and reserve their own judgment on whether its still valuable to them or not!

Medie fra @deedydas
♥ 249↻ 37💬 53🔖 262
https://x.com/deedydas/status/2087037735819460675
LAW
EU AI Act — officiel tekstArtikel 50 om transparens og detekterbart AI-outputhttps://eur-lex.europa.eu/legal-content/EN/TXT/PDF/?uri=OJ:L_202401689

Nyhedsbonus

Gemini har passeret én milliard månedlige brugere og er ifølge Google det hurtigst voksende produkt i selskabets historie. Mere sigende end tallet: 63 procent taler til den, hver femte Live-session bruger kamera eller skærmdeling, og Android-versionen kan handle i over 40 apps. Multimodal agent er blevet distributionsplatform, ikke demo.

1B
Gemini passerer én milliard månedlige brugereVoice, Live-kamera, 150 mio. billeder dagligt og handlinger på tværs af appshttps://blog.google/innovation-and-ai/products/gemini-app/one-billion-monthly-users/