Leading AI companies Anthropic and OpenAI are moving to integrate specialized safety evaluation teams and processes directly into their organizational structures. As large language models (LLMs) evolve at a rapid pace, the industry is grappling with how to balance safety with development speed.
Both companies aim to embed safety evaluation processes across the entire lifecycle of their models, from training to deployment. Traditionally, independent third-party organizations were considered ideal for conducting AI safety assessments. However, as development accelerates, establishing a framework to detect and mitigate risks in real time right on the front lines of development has become essential.
The biggest point of contention is "evaluator independence." By housing evaluation functions internally, how will conflicting incentives—namely, business-driven development goals versus ensuring safety—be reconciled? Maintaining and operating objective evaluation criteria while preserving transparency will serve as a crucial touchstone for future AI governance.