All editions

Two frontier labs are inviting evaluators in. Authorising your agents stays with you.

20 September 2026

Every piece read against the primary source. Where the coverage gets it wrong, the note says so.

Five pieces on pacing the frontier, and on who checks what. A lab chief who invites outside evaluators in, and a rival who promises the same. A playbook that locates the bottleneck at authorisation. A CIO who calls the technology increasingly a commodity. Experts who say European businesses usually prefer the most capable models. And a European strategy that wants pacing verified and clauses against lock-in in procurement contracts. Curated and commented, not aggregated.

Dario Amodei
12 Sept 2026
We Must Pace the Frontier

The piece of the week, and the only one in which a frontier lab commits to something. Amodei argues that capabilities have been growing markedly faster since the summer, mainly because AI is increasingly helping to build the next generation of AI, and that operational excellence, alignment, interpretability and testing need time to catch up. His plan has three steps. In the first, each frontier company gives a team of third-party evaluators "ongoing, employee-like access". Anthropic says it is taking this step unilaterally and will equip its own evaluators with desks, access badges, company laptops and the right to publish findings without editorial control by Anthropic; he calls on governments to require the same of the other frontier companies. In the second, labs in democratic countries coordinate on safety standards and limits on the rate of unchecked progress; for antitrust reasons, that needs government mediation or waivers. In the third, democratic governments try to coordinate with authoritarian ones, as far as that is possible. Sam Altman answered the same day that OpenAI will do the same; Elon Musk wrote "Dario is right". Pacing, in Amodei’s words, "does not mean halting model training or technical progress", and progress will stay "relatively fast". The evaluators are announced for "the near future", without a date, with METR named only as an example. What gets checked is the lab: its safety practices, its models, its training pipelines. How its customers deploy those models does not come up; customers appear once, in the exception that keeps their private information out of the evaluators’ reach. Under the AI Act, providers of general-purpose AI models with systemic risk already owe model evaluations, including adversarial testing, and documentation for those who build on their models. What the essay adds is independence and publication.

World Economic Forum · Capgemini
May 2026
AI Agents in Action: A Playbook for Trusted Adoption, Authorization and Scaling

The deployment side of Amodei’s first step, written four months before it. The World Economic Forum and Capgemini place the bottleneck in agent deployments at authorisation: "organizations grasp what an agent can do, but they struggle to define what it should be authorized to do in context." Their proposal is a written authorisation profile per deployment, with scope, operating context, authority, controls, required evidence, accountable owners and a re-authorisation cadence. Two passages connect it to this week. "Standard model benchmarks alone are insufficient", because they often miss tool misuse, cascading failures and unintended interactions between agents. And because many of a company’s agents may share one foundation model, "a single model-level vulnerability can propagate across an organization’s entire agent estate simultaneously, reinforcing the need for deployment-level authorization and monitoring for each instance". An evaluator at the lab may find a flaw in the model. Authorising the agent that inherits it stays with the company that runs it. In my view a new model version belongs among the triggers for re-authorisation, and only a contract that requires notice of it makes it visible. The playbook is a framework, not a study: it contains no survey and no data, and says of itself that it "does not offer universal answers".

CIO.com
9 Sept 2026
The AI employees are already on the floor. Is anyone watching?

The view from operations. Naren Gangavarapu writes for CIO.com from a deployment at one of Australia’s largest tourism and cruise operators, and for him the real work starts after go-live: tolerance limits, what line managers may override, how an incident is contained, and that every override is logged, because "Override without logging is an invisible governance failure." His conclusion, and the sentence the next piece contradicts: "The technology is, increasingly, a commodity. The governance maturity is the differentiator." That is a claim about the companies deploying AI, not about the labs. Read that way, a new model reaches a company only as fast as the company can re-approve its agents on it. The results from his deployment come without a period, a method or a named company, so they are left out here.

European Commission · AI Office
15 Jul 2026
Enhancing competitiveness, sovereignty and security of the European Union in frontier AI

The counterweight to that sentence comes from an expert forum the European Commission’s AI Office convened in April: more than 100 experts under the Chatham House rule, summarised in July in a report that explicitly does not represent the Commission’s position. On buyers, some of the experts run against Gangavarapu: "European businesses usually prefer the most capable models available", and industries’ purchasing patterns "already suggest a strong preference for frontier over near-frontier systems". The report gives no data for either statement. Both can hold in one company: the agent answering supplier queries runs on whatever is cheapest, the one working on engineering data needs the frontier, and only the second has few providers to switch to. What makes the report useful is a definition one of the experts offers: sovereignty as "the technical capacity to audit and evaluate models on European soil, together with the institutional readiness to switch providers when conditions change". Carried over to a company, those are two questions for any AI contract: can we test what we receive, and can we change provider? The second depends on the first, because a change of provider takes as long as your own evaluations take on the new model, and a fine-tune stays behind. The report also notes that the most capable systems are increasingly deployed internally by their developers before external release. That is where embedded evaluators, who are also to check deployment practices at the lab, would see them first.

transformative-ai.eu
14 Sept 2026
A Transformative AI Strategy for Europe

And the European strategy that appeared two days after the essay; it covers developments only up to early September. The group was convened by Monika Schnitzer, with Daniel Privitera as editorial lead; its Senior Expert Council, which reviewed drafts and gave strategic advice, includes Yoshua Bengio, Daron Acemoglu, Philippe Aghion and Margrethe Vestager, who, the strategy notes, do not necessarily endorse all of it. Its focus is guaranteed access to foreign frontier AI; a European frontier project is covered too, but treated as extremely costly. Two passages meet this week’s question. "Dealing with frontier AI risks might require international agreements on ‘pacing’ AI development, and Europe can help negotiate and verify such treaties", and elsewhere, as one example of why states might pace or pause, "to ensure that external evaluators have enough time for safety testing". And on procurement: "Procurement contracts should include clauses on exportability of data and workflows in order to prevent lock-in or dependence on a single provider." That clause is written for EU institutions, and how AI adoption and diffusion are best supported lies outside the strategy’s scope. The clause carries over anyway. A buyer who can move data and workflows to another model decides when to take the next version. Without that, the lab’s calendar of releases and retirements sets the pace, whatever the speed of the frontier, and so does any decision to cut access.

Related: Agentic AI — whether it pays, build vs buy, and where you stand →

The reading list by email once the mailing starts: five pieces, my commentary. Unsubscribe anytime.