Your AI Agent Scores 80%. Which Tasks Can It Repeat?
A small Python evaluator shows why task-level outcomes and clean trial resets belong in the same agent test harness.
Search for a command to run...
A small Python evaluator shows why task-level outcomes and clean trial resets belong in the same agent test harness.
Meta's stated shift toward loss and reordering tolerance points to the missing contract in many distributed training stacks.
Cursor Origin points at a real bottleneck: once agents create changes cheaply, a forge has to ration reviewer attention with evidence and explicit merge lanes.
LaCache makes the case for exact reuse, but the serving win comes from coordinating cache identity and precision at the same boundary.
A 5-7 MB WASM embedding model changes the tradeoff: for many docs and support flows, the right search layer now lives in the browser, not behind another inference API.
Fusing cheaper models matched Claude Fable 5 at half the cost. OpenRouter's own numbers say about 75% of the gain came from one rewrite pass. Here is how to capture it without paying for six frontier calls.
Open-weight models keep getting cheaper to serve. The reason is memory, and a wave of KV-sharing and compressed-attention tricks is quietly rewriting agent economics.
