Evaluation moves inside the laboratory
Anthropic has announced a partnership with Accenture to test a new model of independent evaluation from inside a frontier AI company. The 18 September arrangement will be led by Faculty, Accenture's specialist AI business, and cover model evaluation, red-teaming, alignment assessment and safeguard testing. Anthropic says embedded evaluators will receive access comparable to employees, allowing them to observe models and decisions during development rather than only testing a finished release.
Earlier access could reveal problems while they are still easier to fix. External evaluations often receive limited time, documentation and model access shortly before launch. An embedded team can follow changes, speak with staff and assess whether internal practices match public commitments. The approach also creates a difficult balance: evaluators need enough integration to understand the work without becoming so dependent on the organisation that they lose critical distance.
Enterprise experience shapes the review
Accenture and Faculty bring experience deploying AI in business and government. Anthropic argues that this practical knowledge can help evaluators test how safeguards behave in real use rather than only in abstract benchmarks. A model may perform well on a standard refusal test yet fail when connected to enterprise tools, mixed with confidential data or given authority across a long-running workflow.
The evaluation scope should therefore include system behaviour, not only model output. Tool permissions, logging, escalation and human oversight influence the risks users experience. At the same time, an evaluator with enterprise clients may have commercial interests connected to deployment. Clear conflict rules and separation between assurance work and sales activity are necessary for the public to interpret findings.
The investment is large, but funding is complicated
Anthropic and Accenture each expect to invest at least US$1 billion over five years in capacity connected to this area. The announcement says Anthropic will directly fund Accenture's evaluation work because no pooled or government mechanism currently exists. Anthropic also says it is discussing separately funded pilots with nonprofit evaluators including METR.
Direct funding does not automatically invalidate an evaluation, but it creates a structural conflict that must be managed. Contracts should protect the evaluator's ability to report serious findings, retain evidence and publish limitations. Fees should not depend on a favourable result or a model release proceeding. Long-term, shared funding and common accreditation could reduce dependence on the company being examined.
Standards and reporting are still open questions
Anthropic acknowledges that no settled standard defines what embedded evaluators should access or how they should report. Employee-like access may expose trade secrets, personal data and security-sensitive information. The evaluator needs secure handling rules and enough freedom to describe material risks without revealing details that enable misuse. An external oversight process may be needed when the parties disagree about disclosure.
Reports should state the scope, methods, access constraints, unresolved findings and the developer's response. A general statement that a model was tested is not sufficient. The public also needs to know whether the evaluator reviewed the final deployed system or an earlier checkpoint. Repeat work after material changes will matter because safeguards, tools and product settings can alter risk after the initial assessment.
Accountability remains with the developer
Customers can reinforce the model by asking for the underlying evaluation scope during procurement. They should distinguish access to the base model from testing of the final product, which may add tools, memory and different system instructions. Assurance is strongest when it follows the system that users actually encounter and is repeated after material changes.
A credible program should also include protections for evaluators. Staff need channels to raise concerns outside the immediate project, protection from commercial retaliation and rules for handling disagreements about severity. Rotating personnel can reduce capture but may weaken the context that makes embedding valuable, so governance must balance continuity with independence. Regulators and civil-society experts could advise on reporting templates without receiving confidential model details. Over time, comparable reports across laboratories would help customers distinguish rigorous evaluation from a bespoke engagement whose conclusions cannot be compared. The partnership can contribute most by publishing methods that others can challenge and improve.
Anthropic explicitly says embedded evaluation does not transfer responsibility for model safety. That is an important boundary. An evaluator provides evidence and challenge; the developer decides whether to train, release or restrict a system. Regulators and customers should not treat third-party involvement as a blanket certification, particularly while methods and standards are experimental.
The partnership is a serious attempt to make frontier development more observable. Its value will depend on evaluator independence, access to consequential decisions and transparent reporting of both findings and constraints. Multiple evaluators with different methods would provide stronger assurance than one embedded partner. If Anthropic and Accenture publish concrete lessons, including disagreements and changes made, the experiment could help establish a credible profession around continuous frontier-model evaluation.