Answer in brief
My August piece read one lab containing its own research environment as the signal. That still holds and it is no longer the whole picture. Since then one frontier lab has proposed giving outside evaluators badges, laptops and the right to publish what they find, and another has asked the defensive side of the industry to move together. The argument went from what I secure, to what someone else can check, to what we do jointly.
What I said in August
Briefly: a lab paused some frontier training and tightened isolation, logging and monitoring around its own research environment. What mattered for the rest of us was where the trust boundary moved. Once a model can run code, call tools or reach a network, it is an actor inside your environment rather than an artifact under test.
I stand by that. What I would add is that it is a claim a lab makes about itself.
The second layer: someone else has to be able to check
Dario Amodei's September essay proposes pacing the rate of capability advance, and is careful about what that means:
"pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this."
The first of his three steps is the one Anthropic commits to on its own: embedded third-party evaluators with "ongoing, employee-like access," verifying that a company follows the practices it published. He describes inviting a review team with "desks in our offices, access badges, and company laptops."
The clause I keep returning to is what reviewers may say afterward. They may publish findings "without editorial control," the company may redact narrow categories, and then:
"The reviewers can say publicly if a redaction removed something important to their conclusions."
That sentence is what makes the rest checkable. A redaction right with no disclosure is a hole. One where the reviewer can say a hole exists is a signal. That is the difference between a control you assert and a control someone else can inspect.
The small-scale version, from my own operation
I run a fleet of AI agents on real work and I have written rules for how they behave. Writing the rule was never the part that worked. My operating file carries one that makes the agents check each other:
"Before you act on (a) a negative or uniform result, (b) a correction across a document family, (c) anything that becomes a card for Chris, or (d) after ninety minutes of solo work, state the claim to your partner lane and wait for a contradicting fact."
Here is what it bought on one ordinary evening. A coordinating agent relayed a measured fact: the calendar still did not carry a client visit. The receiving agent did not take it. The record reads:
"the lane re-read the calendar before writing it and found Chris had blocked the slot minutes earlier. A measured claim carried forward by relay is a memory, not a measurement."
Nobody was careless. The measurement was true when it was taken. It was caught because checking it was somebody else's job, which is the same shape as an evaluator with a badge.
The third layer: you cannot do this one alone
The newest OpenAI piece is not about one lab's environment. It calls for defenders to move together, and its governing sentence is explicit about who is included:
"Each of us can reduce risk now. All organizations, cybersecurity companies, technology partners, governments, and AI frontier companies have an important role: accelerate defenders' priorities with tools, funding, and hands-on support, especially for critical infrastructure organizations with limited budgets."
Its framing of why now is a window rather than a deadline: "Today's AI advances are already giving defenders new ways to fix weaknesses that have accumulated for years." Hundreds of organizations had signed when I captured it, though that is a snapshot.
What I take from the three together
The order matters, and it is not the order most organizations follow. Contain your environment, because that is yours to do. Then arrange for someone outside to check, and to say so if they were prevented. Then find who else is doing the same, because a defensive floor is not something one organization stands on alone.
Most AI governance work I see stops at the first step, because that is the step you can finish and show a board. The second turns a claim into evidence, and it is uncomfortable, because it means someone can publish that you fell short.
I do not think the labs will get all of this right, and none of these commitments has been tested yet. The direction of travel is the useful part, and it points somewhere ordinary organizations can follow at their own scale. The question worth asking about your own AI governance: who, other than the team that built it, can check that it works, and are they free to say what they found?