Dagens Vibes — 22. august 2026
Modellen er kun motoren. Ox Alpha ligner en model fra GLM-5.X-familien og får sin første ordentlige prøvetur; NVIDIAs harness klarer alle 183 offentlige ARC-AGI-3-niveauer; og Pi spørger høfligt, om MCP overhovedet behøvede komme med.
Fra X-feedet
Ox Alpha overtog feedet: et gratis stealth-drop med 1M context og multimodale input. Plinys agent peger på Z.ai og GLM-5.X-familien. Det er ikke officielt bekræftet, men et væsentligt spor. Robin Ebers’ 20 videoredigeringer giver det praktiske vibe check.
A mysterious new AI model just appeared. Ox Alpha offers a 1M context window, multimodal capabilities, zero data retention, and nearly unlimited usage for an entire week. OpenCode says it has capacity for 100 trillion tokens per day. That’s 1.16b tokens per second. Where the heck did they get that much compute? Nobody knows which company built it.
Ox-alpha is from Zai, GLM-5.X family my agent has spoken 🙌 https://t.co/VT22xeuVkX
Ox Alpha is a powerful model i've put it into my new AI video editor that editor currently runs on Fable 5 @ Low and it... was fucking impressive gave it two youtube videos (23 min and 28 min) each video was edited 10 times 1) on the clean video it landed Opus 5 @ High's cut almost word for word (7 of 9 finished runs were 90–97% the same cut) 2) it edits for substance unprompted. deduped retakes down to the last good take, something Fable needs to be told to do. 3) ~7x faster than Opus 5 @ High on the clean video. just a few problems... - 4 of 20 runs broke: 1 invalid answer, 3 thought forever and never answered (all 3 on the harder video) - it gets scattered on messy footage (runs agree 78% with each other there vs 88% on clean) - we still don't know what it will cost so one thing i can say for certain: on clean footage it edits at Opus 5 @ High level, which no model ever did. hard to ignore. in a single sentence, on easier problems, the model outperforms Opus 5 @ High, but when things get messier, it falls behind fast Ox Alpha is Opus's taste/insticts, but unfortunately without Opus's discipline. BUT it is a lot faster!
DeepSeek svarede med en officiel multimodal Flash-model til API’et. Den matcher ifølge DeepSeek tekstmodellen og nærmer sig Opus 4.8 på visuelle agentbenchmarks; billeder koster højst 384 tokens ved almindelig V4-Flash-pris. Ingen solid uafhængig vibe check i feedet endnu, så stjernerne bliver siddende på producentens uniform.
DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model. 1/n
NVIDIAs AVO løste alle 183 niveauer i ARC-AGI-3’s offentlige sæt med Opus 5, hvor standard-harnesset lå omkring 30 procent. Det er ikke et privat leaderboard-resultat eller en ren 30→100-ablation, men persistent hukommelse, feedback og en supervisor flyttede systemet voldsomt. Samme loop har tidligere kørt syv døgn på GPU-kerneloptimering. Harness engineering er blevet modelarbejde uden vægtløftning.
ARC-AGI-3 is solved. AGI is here. It’s not a debate. Inflammatory claims aside, NVIDIA solved the ARC-AGI-3 public benchmark with their agent harness, AVO + Opus 5. Opus 5 + AVO boosted Opus 5's score over the standard ARC-AGI harness from 30% --> 100%! It's what a great harness does: exposes latent capabilities of an LLM. For AVO it's: 1. Persistent memory across context windows 2. An inspect → plan → implement → evaluate loop 3. Execution feedback so the agent learns from mistakes at inference time 4. A supervisor that detects stagnation in failed trajectories and redirects the agent 5. External statefulness of the conversation transcript - tool calls, execution artifacts and the agent trajectory The great thing about frontier harness engineering is that you don't need to work at a frontier lab to contribute. Harness advances don't necessarily need to touch the model weights, but can make all the difference.
Et studie kørte en softwareopgave gennem syv agenter og fem modeller: i et domæne med moden CLI løste agenter uden indbygget MCP opgaven lige så stabilt og 5–28 gange billigere. Det er ét afgrænset setup, ikke MCP’s dødsattest — men “brug de eksisterende kommandolinjeværktøjer” er et mistænkeligt effektivt plot twist.
This week we read research from a team of academics that ran a software task across 7 agents and 5 models. They found that in domains with a mature CLI ecosystem, agents without MCP baked in completed the task just as reliably and were 5-28x cheaper. Full arXiv paper below https://t.co/DOtkMeqpoC
To små arbejdsformer med høj brugsfaktor: Matt Pococks /implement-spec splitter en spec ud i parallelle worktrees og reviewer samlet mod originalen; Arnav Gupta peger på medium effort og Pi’s lille systemprompt som sweet spot. Maksimal tænkning er ikke altid maksimal intelligens. Nogle gange er det bare dyrere maskinbekymring.
I'm trying out an /implement-spec skill Essentially a multi-agent implementer that: - Takes in a spec and tickets - Does codebase research in a subagent - Implements all the tickets in subagents with maximum concurrency - Reviews the final code against the spec - Cleans up all worktrees Should be able to smash out huge chunks of work autonomously with minimal supervision. https://t.co/lTmPYXkUx7
I've been using medium forever. When I show others people have a range of emotions from surprised to dismissal. Most people have so much AI psychosis they dont care to evaluate models and harnesses properly before figuring out the sweet spot for their work. For example Pi visibly gets tasks done faster than Claude Code or Codex. You don't need benchmarks to prove it (Databricks proved it if you want though), simply because Pi runs on a much smaller default context size because of smaller system prompt and less tools. Similarly you don't need benchmarks to prove this but in medium, the model runs fastest, gets tasks done sooner and doesn't entangle itself in its own bullshit. But if you need it there is a benchmark here (thanks @cline ) Medium > xhigh/max not just for Opus but also for Sol.
Nyhedsbonus
Anthropic åbner Mythos 5’s offensive styrke gennem en smal defensiv luge: Claude Security kan scanne ejede repositories og returnerer sårbarheder, CWE, severity og foreslåede patches — men brugeren får artefaktet, ikke fri adgang til modellen. Samtidig lægger de 35 millioner dollar i credits til open source-sikkerhed. Capability containment som produktdesign, ikke bare en større advarselsboks.
DeepMind gør EVE Online til næste agentlaboratorium: en verden med 23 års økonomi, alliancer og konflikter skal teste hukommelse på tværs af context windows, planlægning over måneder og multi-agent-forhandling. Første stop er en offline sandbox — robotterne får ikke adgang til rumkapitalismen uden prøveperiode. Fornuftigt.