A model lab sets an unusual evaluation policy
TypeSafe AI has outlined a model-release policy that deliberately steps away from the familiar wall of public benchmark scores. In a post dated 11 September, the company says it will not use standard benchmark tables for its releases. Instead, it plans to publish dated evaluation snapshots, retire them rather than repeatedly optimise against them and share its evolving internal evaluations with their limitations.
The announcement is more than a general complaint about leaderboards. It defines how the company intends to support claims for System One Models, including Jev. TypeSafe says its new model category does not map neatly onto conventional language-model benchmarks, but it also argues that repeated exposure turns any public test into a target that influences model selection and training decisions.
Why repeated testing can distort results
The post uses the term “benchmaxxing” for the process by which teams improve on a known evaluation, even without directly training on its questions. Researchers can try many model variants, prompts and post-training settings, then keep the version that scores best. The benchmark has shaped the result through selection, so a high score may reveal less about performance on unseen work than the headline suggests.
TypeSafe also points to a feedback loop between rankings and market expectations. When a result clashes with what users believe about a model, an evaluation may be revised or replaced until it appears more plausible. Such revisions can be defensible, but they weaken the idea that a public number is an independent and stable measurement of general capability.
What TypeSafe says it will publish
Under the stated policy, new evaluations will be dated snapshots rather than permanent targets. TypeSafe says it will disclose caveats, possible cherry-picking and evidence that makes its models look worse, while de-emphasising benchmark victories even when it is ahead. Its launch material later applied this approach by publishing workflow tests alongside qualifications about authorship, reference models and the likelihood that headline gains represented favourable cases.
This approach does not eliminate conflicts of interest. The company still designs many of its own evaluations and decides what to publish. A dated test can still be selected selectively, and retiring it can make longitudinal comparison harder. The value of the policy will depend on whether TypeSafe preserves the underlying methods and results, reports failures consistently and resists replacing an awkward measure with a friendlier one.
The burden shifts towards customer evaluation
For buyers, the practical message is to build private tests from real workflows. TypeSafe argues that its structured System One tasks are easier to evaluate than open-ended prose because the possible decisions are defined in advance. That can support clear measurements of accuracy, calibration, latency and abstention, but only if the dataset reflects the organisation’s actual cases rather than a polished demonstration set.
A useful evaluation should be held back from prompt and model tuning, include rare and ambiguous examples and measure performance across confidence bands. Teams should compare the operational result, including retries and review time, rather than a single accuracy figure. If probabilities are central to an automated decision, calibration drift should be monitored after deployment as inputs and business rules change.
Transparency will determine whether the policy works
TypeSafe’s position is a useful challenge to the model industry’s reliance on familiar leaderboards, but it is not a substitute for evidence. Customers still need enough reproducible information to compare releases, understand regressions and decide whether a new model is safer or more economical. Private evaluation helps the customer; transparent shared evaluation helps the wider market detect exaggerated claims.
The policy will become meaningful through repetition. Each future release will test whether TypeSafe consistently dates results, states what changed, exposes unfavourable outcomes and explains when an old evaluation is no longer used. If it does, the company could provide a more honest record than a constantly rising benchmark table. If it does not, “no public benchmarks” could simply reduce outside scrutiny. The September commitment gives customers a standard against which to judge those future releases.
There is also a governance question for customers that use vendor evidence in procurement. A score copied into a business case can outlive the model, harness and dataset that produced it. Buyers should record the date, model version, evaluation code and decision supported by each result. They should distinguish vendor demonstrations from independently reproduced tests and their own acceptance criteria. That record makes it possible to explain why a model was selected and when it must be reassessed. TypeSafe’s snapshot idea is most useful when it encourages that discipline rather than allowing old evidence to disappear from view. Public archives would strengthen the approach by letting researchers compare successive snapshots without turning one frozen test into the permanent target. Clear change logs could show why an evaluation was retired and what replaced it.