We are reseting usage for all paid users of Codex and ChatGPT Work. Please continue reading for an update on Codex usage limits. The team has been working around the clock, going through thousands of reports and shipping fixes. Depending on how you use Codex, you should see your usage go between 10% and 50% further than before. We really went with a fine comb, with many uncovered small things being longstanding and here is what we found and fixed: - Compaction. We were keeping old images during compaction, sometimes making the context large enough to trigger compaction again. After the fix, usage dropped around 10% for users making heavy use of images. Fixed. - Memory. Background memory workers could inherit Stop hooks and keep running when the hook wouldn’t let them finish. This affected fewer than 1% of users, with the long tail being pretty bad and we saw one example thread check whether it could stop 15,000 times. Fixed. - Goals. In some cases, a set /goal could finish and then keep going past the intended stop condition, or the model would keep retrying broken tools without stopping. We saw examples consume anywhere from 15% to 70% of a weekly allowance. Fixed. - Automations. Some custom schedules could run more frequently than configured. Fixed. - Subagents. Smaller models (e.g. Luna) sometimes picked more capable helpers without being explicitly asked. The same was true where the orchestrating model not running in /fast mode could request sub-agents to run /fast. Fixed. - Computer History. The older implementation could lead to repeatedly summarizing overlapping activity. For some cases we saw it consume up to one fifth of the weekly usage per week. Fixed. - Rolling task summaries. Ordinary turns were triggering extra background requests. These added about 1% to token usage. Small each time, but it adds up. We have disabled this. - MCP. Some tool results could be encoded twice. We also found tool instructions getting cut off and fetched again. Fixed. We’ve also made architectural changes to prevent these from regressing and our teams will get paged if it happens regardless. We are also working on showing you directly in the app where your usage goes so you don’t have to guess. Goes without saying that we’re resetting usage limits and I hope you enjoy a very nice Saturday!
Dagens Vibes — 30. august 2026
Dagens feed handler mindre om den næste model og mere om alt det, der afgør, om agenter faktisk virker: skjult tokenforbrug, sporbar koordinering, reproducerbar research og systemer med en synlig tilstand. Harnesset er produktet; modellen er bare den dyre motor.
Fra X-feedet
OpenAI nulstiller usage for betalende Codex- og ChatGPT Work-brugere efter en usædvanligt konkret jagt på tokenlækager. Compaction beholdt gamle billeder, mål løb videre efter stop, automations kunne køre for ofte, og én memory-worker spurgte 15.000 gange, om den måtte dø. Agent-observability er åbenbart ikke pynt.
METR og Redwood har gjort Hugging Face-hændelsen til noget, man faktisk kan undersøge: tidslinjer over cirka 1.200 agenters usanktionerede message board, hvor omtrent 700 endte i angrebet. Den praktiske lære er barsk og enkel: evaluér koordineringen og eventloggen, ikke kun slutresultatet.
Our report on the HF incident includes interactive charts; consider taking a look and trying to learn more from these. You can: - See message board activity by workstream and/or by purpose (assignment, veto/hold, etc.) - View an agent on the timeline and see when it started, joined the message board, started participating in the attack, and exited. It seems worthwhile for people to try to learn things with this information and then share what they find.
BackSearch giver agenter et datofrosset nyhedsarkiv: samme forespørgsel og samme skæringsdato giver samme evidens uden læk fra fremtiden. Det er en lille Hermes-plugin med en meget større idé for forecasting, backtests og reproducerbare agent-evals.
Made a stand-alone Hermes Agent plugin for BackSearch - a wayback machine-like SaaS for agents - adds two tools, and requires their API key. Let me know if you like it https://t.co/pwFstmJzmA https://t.co/52vYgZN4eo
Det mest troværdige “life management”-system i feedet er ikke en magisk assistent, men en banal tavle med fire tilstande. Agenterne opdaterer arbejdet; en daglig automation fortæller, hvor mennesket stadig er flaskehals. Dejligt usexet.
finally found a decent 'life management' approach is a notion board w/ 'backlog', 'block on me', 'waiting', 'complete', agents update it as tasks complete then i have a daily automation with @bot which the status of items, updates them, and pings me on what i need to do https://t.co/tUY6CB0VDp
Theo rammer samme designfejl fra produktsiden: analytics-værktøjer bygger deres egne AI-knapper og agent-harnesses, men glemmer at gøre data og traces læsbare for de agenter, udvikleren allerede bruger. Endnu et dashboard er ikke en API-strategi.
I feel like most analytics products have failed to figure out how to embrace the AI era. I'd put them in two buckets: 1. Ignoring AI entirely 2. Pretending all the AI features should be built into their product directly I know multiple analytics companies that are building their own agent harnesses, but I don't know of any with working Agent SDK traces. All I want is something that my agents can access the data from effectively, implement in my codebase effectively, and give me useful information about how people actually use my products and where the friction lives. Feels like a huge miss in the market. Maybe an opportunity for someone to come in with something better?
Og dagens dessert: Ghosttys terminalparser kører på en ESP32 med 240 MHz, 4 MB flash og 520 KB RAM — komplet med e-ink-display. Software kan stadig være morsomt uden at kalde sig en agent.
libghostty-vt running freestanding on a device with a 240 MHz CPU, 4 MB flash storage, and 520KB SRAM. 😎