Apollo Research proposes principles for embedded evaluations; implementation decides their value
AI safety evaluator Apollo Research published principles for embedded evaluations, proposing claim-by-claim verification of pre-fixed safety claims; developer adoption remains to be seen.
Apollo Research, an AI safety evaluation organization, published its principles for embedded evaluations on September 30, arguing their impact depends on evaluator access, resources, and the weight findings carry in real decisions.
The core proposal is claim-based assessment: instead of an overall judgment of the developer, evaluators verify specific claims fixed in advance, such as "Models never attempted to disable or evade their monitoring during internal deployment." Each claim ends with one of five verdicts, insufficient access scores worst, a problem the developer reports itself scores better than one the evaluator finds, and the developer cannot veto the verdict.
The proposal also calls for public reports by default, with the evaluator's conclusion never redactable, and argues voluntary commitments are unlikely to suffice, so embedded evaluations should eventually be required by law. This is a unilateral proposal by the evaluator; no developer has yet said it will adopt it.
Sources:https://www.apolloresearch.ai/blog/principles-for-embedded-evaluationshttps://www.lesswrong.com/posts/jD32DPYdZEcyLZqto/principles-for-embedded-evaluations