Skip to content
wildcard

InsightsAI news5 min read

AI weekly: cheaper models, harness evals, week of 2026-10-09

Haiku 5.5 moves the small-model price floor, OpenAI's Cookbook shows why tool tests lie, a US open model arrives. What it means for a DACH engineering leader.

Cliff des Ligneris

This week’s news is about the bill, not the model. A small model now costs a tenth of what it did a year ago, and the eval recipe OpenAI published shows that your green tool tests say nothing about the agent your users meet. If I ran engineering at a 100-person DACH product company, I would spend Monday on two things: a hard spend cap on every API key an agent can reach, and one eval that runs my MCP tools inside the real harness.

Top highlights at a glance

  • Claude Haiku 5.5 lands at $0.10 per million input tokens and $0.50 output. Anthropic says it costs about 75% less to run than Haiku 4.5 and ships with an adjustable effort setting.
  • OpenAI’s Cookbook published a harness-aware eval recipe from deepsense.ai. The same MCP tools pass in a bare function-calling loop and fail inside Codex.
  • Reflection announced Beam, a 501B-total, 23B-active open-weight model trained from scratch in the US, with Apache 2.0 weights promised this month.

Second briefing in the series. Last week set three baselines; here is how each one moved.

  1. Cost engineering is a product skill. Continuing, one tier down. Last week the price step came from the near-frontier tier: GPT-6.1 Sol at a fifth of Astra, Sonnet 5.5 up to 30% cheaper. This week it reaches the small-model tier. Haiku 5.5 matches GPT-6 Luna’s price up to 100k tokens, Sonnet 5.5 cache reads were halved, and OpenAI reports LegalOn cut daily Codex cost by 65% by routing tasks to different models. Which model runs which task is a weekly decision now.
  2. The harness is the product. Continuing, with the focus shifting. Last week harnesses stabilised: Pi 1.0, sandbox examples per cloud. This week the question is how to run and evaluate them. Two Kubernetes creators are building a cloud-native harness, and the Cookbook recipe shows plugin behaviour changing when the same tools run inside a product harness. Teams that only eval the model will be surprised in production.
  3. Open weights meet the cyber-risk argument. New this week. Beam arrives as a US open model, Anthropic expands its Cyber Verification Program, and Nathan Lambert argues the open-weights debate is stuck in a lose-lose. Expect sharper procurement questions about open models.

Last week’s third baseline, narrow models for narrow decisions, has not gone away. It shows up in Haiku 5.5’s effort setting and in LegalOn’s task routing, so I have folded it into the cost theme.

Models, harnesses and technical frameworks

Claude Haiku 5.5

Source: https://www.anthropic.com/claude-haiku-5-5 and https://simonwillison.net/2026/Oct/7/claude-haiku-5-5/. Category: Models.

Anthropic released Haiku 5.5 on 7 October at $0.10/$0.50 per million tokens up to 100k context, and $0.50/$2.50 beyond. Anthropic’s own table puts it at 39.2% on Terminal-Bench 4.0 against 70.6% for Sonnet 5.5. It is the first Haiku with an effort setting. Simon Willison measured the new tokenizer at about 1.25x the token count of Haiku 4.5 on the same long prompt, which is a hidden price increase.

Why it matters: subagent, classification and summarisation work can move to a model that costs a fraction of Sonnet. Re-run your cost model with the tokenizer change before you quote the savings to anyone.

Harness-aware plugin evaluation

Source: https://github.com/openai/openai-cookbook/commit/5bd6cb40b6b84b69b2032d810610aa85064102df. Category: Evals.

The recipe evals one MCP plugin at three levels: direct tool tests, a standalone function-calling loop, and the Codex product harness. In the recorded run, two of five cases pass in the standalone loop and fail under Codex, because nothing told Codex to ask when two candidates tie. The fix was guidance in the MCP server’s instructions field, not a prompt change in the client.

Why it matters: tool tests can be green while the shipped agent does something else. A level-three eval belongs in CI for every MCP server you expose.

Stacklok’s Mecatl, a cloud-native agent harness

Source: https://www.latent.space/p/stacklok. Category: Agent harnesses.

Kubernetes co-creators Craig McLuckie and Joe Beda are building Mecatl, an open-source harness started in June. It treats agent sessions like orchestrated workloads: isolated tool execution, reliable sessions, preserved context. Their two questions are worth keeping: why is the agent loop coupled to the tool-calling subsystem, and why is agent identity reasoned about through human identity systems.

Why it matters: if your agents run on developer laptops, the code and context they touch are unmanaged. The 2014 container argument is about to replay for agents.

Reflection Beam, 501B-A23B open model

Source: https://www.latent.space/p/ainews-reflection-beam-501b-a23b. Category: Models.

Reflection announced Beam, a text-only mixture-of-experts model with 501B total and 23B active parameters, trained from scratch, with Apache 2.0 weights due this month. Reflection claims 80.9 on SWE-bench Verified. Independent observers quoted in the same roundup place it around GLM-5.2 and below DeepSeek V4.1 Flash on some benchmarks.

Why it matters: a permissively licensed, US-trained model at this scale changes the procurement conversation for teams that cannot send data to Chinese-origin weights. Wait for the weights and the tech report before you plan on it.

Product and engineering strategy

Default hard budget caps

Source: https://simonwillison.net/2026/Oct/3/default-hard-budget-caps/.

Agents remove the friction of spinning up things that cost money, so pay-by-usage services need hard caps by default. AWS shipped project spend limits on 16 September and Google Cloud added Spend Caps in July. Key takeaway: put a hard cap on every API key an agent can reach, this sprint.

LegalOn cuts Codex cost by matching models to tasks

Source: https://openai.com/index/legalon-halves-codex-costs.

OpenAI reports that LegalOn cut estimated daily Codex cost by 65% at unchanged development speed by assigning Astra, Sol and Luna to different task classes. It is a vendor number. The method is the lesson: model routing is a budget line your team owns.

Back to insights

Talk to us.

Thirty minutes. We talk about your situation and you leave with a next step or an honest no.