Dagens Vibes · 2. juli 2026

Dagens Vibes — 2. juli 2026

Dagens hovedvibe: agent-æraen er ved at blive infrastruktur. Fresh machines, hot-reloadede skills, eval-flywheels, kinesiske open models og finansfolk der pakker OpenAI-stakes ind som sikkerhed. Sommer, men med datacenter-sved.

Fra X-feedet

Feedet pegede mindre på “hvilken model er smartest?” og mere på arbejdsformen rundt om modellerne: hvor agents kører, hvordan de lærer værktøjer, hvordan man måler dem, og hvem der betaler strømregningen.

Kodeforståelse er ikke blevet gammeldags — det er blevet en ny kernekompetence. Geoffrey Litt rammer dagens bedste anti-magiske pointe: når agents skriver mere kode, skal vi være bedre til at forstå, styre og verificere den.

Geoffrey Litt
@geoffreylitt

Hot take: I think it's still important to understand the code that our agents write! In this mega thread (based on my AIE talk today), I will explain why that's the case, and show some ideas for how to efficiently understand code. Alright, let's dive in. 1/ https://t.co/765DNZh6LN

Tweet media
♥ 43 · ↻ 4 · 💬 4 · 🔖 29
https://x.com/geoffreylitt/status/2072522251300409556

Amp-orbs er dagens tydeligste skifte i devmiljøet: agents flytter fra “min maskine” til ephemeral arbejdsrum, lidt som buildservere blev til CI. Det er kedeligt-infra på den farlige måde.

Thorsten Ball
@thorstenball

We launched agents in orbs yesterday. I truly believe that what we think of as a development environment will change dramatically in the next year. These models are incredibly good now: they need less handholding, they need less oversight; inputs in the form of codebase & prompt determine the quality of the output more than oversight; they are exceptionally good at building their own tools now. You can literally throw them on a fresh computer, give them the code and the prompt and they'll figure out a lot on their. And you can spawn an infinite number of them. And they can work in parallel. And they are patient, they wait for you. Combine all of these forces and it starts to make less and less sense to keep agents running on a single, bespoke machine. What we'll see, I think, is similar to the move from a single, pet-like build server to CI systems with ephemeral VMs. How exactly we'll get there and what we'll find along the way? I don't know and I don't think anybody knows. But I'm very excited about finding out.

Referenced: Thorsten Ball @thorstenball

You can now spawn Amp agents in orbs https://t.co/aExDy4YPKt https://t.co/y2vFPU3job

Referenced tweet media
https://x.com/thorstenball/status/2071973919418945860
♥ 74 · ↻ 3 · 💬 7 · 🔖 13
https://x.com/thorstenball/status/2072232955787788547

OpenCode 2.0 går efter hot-reloadede skills, så agenten kan lære nye værktøjer uden at smadre cacheøkonomien. Meget lille detalje, meget stor hverdagsbetydning.

dax
@thdxr

one of the reasons OpenCode 2.0 took so long was we redesigned it for hotreloading if you ask it to make a skill for itself (or make one manually) it'll get picked up immediately in a way that does not bust cache targeting public beta end of week, wish us luck! https://t.co/3h5UQQAwo9

Tweet media
♥ 2.3k · ↻ 54 · 💬 111 · 🔖 188
https://x.com/thdxr/status/2072392464321622452

Matt Pococks /wizard er præcis den slags agent UX der føles som fremtid: ikke endnu en chatboble, men en interaktiv CLI der guider den kedelige opsætning af services.

Matt Pocock
@mattpocockuk

This is just outrageously useful It just had me set up the infra for my personal podcast in about 3 minutes: - Provision a Gemini API key - Set up a Vercel deployment with blob storage - Connect to my podcast app (!) Even had built-in fallbacks for known errors. Incredible

Referenced: Matt Pocock @mattpocockuk

Getting sick of setting up third-party services So I built a skill for it /wizard builds you an interactive CLI for the task you're currently doing, and takes as much work off your hands as possible #1 is how the agent described the wizard, #2-3 is what it looks like: https://t.co/SwMnQOeqX8

Referenced tweet media
https://x.com/mattpocockuk/status/2072042214188847178
♥ 590 · ↻ 12 · 💬 20 · 🔖 745
https://x.com/mattpocockuk/status/2072247137379742174

LongCat-2.0 var dagens “vent, hvad?”: Meituan/open model, 1.6T MoE, MIT, agentic coding, og ifølge tweetet trænet uden Nvidia. Compute-moats har fået en kat i maskinrummet.

Sudo su
@sudoingX

this is the one people should be paying attention to, @Meituan_LongCat just open sourced longcat-2.0, a 1.6T param moe, ~48b active, 1M context, MIT license with weights dropping soon the part that actually matters is meituan says they trained the entire thing on domestic chinese chips, a ~50k card cluster, no nvidia anywhere in the loop and this is the "Owl Alpha" that's been quietly sitting near the top of openrouter for ~2 months pushing ~559B tokens a day, nobody knew whose it was look at where it lands on agentic coding, terminal-bench, swe-bench pro, multilingual, forte, rwsearch, browsecomp, it's sitting right in the frontier cluster next to gpt-5.5, gemini 3.1 pro and opus 4.8, edging gpt-5.5 on swe-bench pro (59.5 to 58.6) it doesn't top every column and it doesn't need to, a delivery company hit near-frontier agentic coding with zero western silicon and then put it under MIT, the nvidia moat was supposed to be the whole game

Tweet media
Referenced: Meituan LongCat @Meituan_LongCat

Introducing LongCat-2.0 🐱 1.6T parameters · MoE with ~48B active · 1M context The full model behind Owl Alpha on @OpenRouter — now available. Built for agentic coding from the ground up: ◆ LongCat Sparse Attention (LSA) — scales efficiently for 1M-context tokens ◆ Zero-Compute Experts — dynamic activation 33B–56B per token, zero wasted compute ◆ MOPD — three specialized expert groups (Agent / Reasoning / Interaction), gate-routed per task How it stacks up: → Terminal-Bench 2.1: 70.8 → SWE-bench Pro: 59.5 (GPT-5.5: 58.6) → SWE-bench Multilingual: 77.3 → FORTE: 73.2 · RWSearch: 78.8 · BrowseComp: 79.9 📖 Tech Blog: https://t.co/4KrjyKiDBn Try it across different scenarios 🧵👇

Referenced tweet media
https://x.com/Meituan_LongCat/status/2071783587205308721
♥ 58 · ↻ 6 · 💬 3 · 🔖 13
https://x.com/sudoingX/status/2072310159599456667

Shopify-casen er en god modgift mod model-hype: evals, repair-loops og promptkompression tog en GraphQL-agent fra $27M til $1M årlig serving cost. Ikke sexet. Derfor interessant.

Shopify Engineering
@ShopifyEng

Shopify's LLMs beat frontier models on a range of tasks at a fraction of the cost. The reason: we put systems in place that enable them to improve themselves, learning from a range of commerce tasks every day. We're presenting our Model Optimization Flywheel at @ICMLconf: a continuous pipeline that turns Shopify's product expertise into robust evals, mines low-scoring conversations, critiques them, repairs them, and feeds them back into the model. Then we compress the prompts without losing quality, so we can make it faster and cheaper. We present an example of the flywheel working at scale: our GraphQL agent. Serving cost dropped from $27M to $1M annualized (−96%). We compressed our system prompt 4× and still beat frontier models on quality. @Drewch and @cmazzaanthony will share concrete recipes, quality-cost-latency trade-offs, and a blueprint you can actually build from. 📅 Monday, July 6 · 11:30am–12:30pm KST 📍 COEX, Hall D1 Link in thread. 👇

♥ 74 · ↻ 6 · 💬 3 · 🔖 58
https://x.com/ShopifyEng/status/2072405411756724677

Remote Labor Index gør Fable-snakken mere konkret: 16,1% af professionelle remote-work projekter accepteret som brugbare. Stadig lavt — men det er hele pointen. Kurven har tænder.

Chubby♨️
@kimmonismus

This is crazier than you might think: Fable-5 now scores 16.10% on the Remote Labor Index What is RLI? The Remote Labor Index uses 240 real remote-work projects from professional freelancers, covering 23 domains and more than $140,000 of human work. Each task comes with the actual brief, files, and accepted human deliverable. Reviewers then compare the AI output against the human reference and ask whether a reasonable client would accept it. That is why the scores are still low. Full projects require planning, file handling, quality control, visual consistency, domain judgment, and final packaging. Fable-5 now leads the public leaderboard at 16.10%. And it’s a crazy jump. We are still deep in exponential development, and now even the toughest benchmarks are being tackled.

Tweet media
Referenced: Center for AI Safety @CAIS

New Remote Labor Index results: AI automation of real remote work is increasing fast. Claude Fable 5 now completes 16.1% of projects at a professional standard, roughly double the next model and up from Opus 4.6’s 4.2% automation rate. https://t.co/juqG3pQcuu

Referenced tweet media
https://x.com/CAIS/status/2072360965522489789
♥ 1.5k · ↻ 112 · 💬 46 · 🔖 536
https://x.com/kimmonismus/status/2072376968729817531

Dagens slop-note: X testede at fjerne de 30 højest betalte revenue-share konti fra For You, og både time spent og DAU gik op. Internettet hvisker: færre incitamenter til skrammel, mere liv. Radikalt koncept.

Nikita Bier
@nikitabier

In a 3% experiment, removing the Top-30 highest paid revenue share accounts from the For You timeline increased both time spent and daily active users on X.

♥ 22.7k · ↻ 784 · 💬 2.5k · 🔖 1.4k
https://x.com/nikitabier/status/2072203879479910490

Nyhedsbonus

Udenfor feedet var nyhedsbilledet samme maskinrum: penge, compute og sikkerhedspolitik. Ikke glitrende, men det er her tempoet bliver bestemt.