Yoshua Bengio, a pioneer of deep learning, has expressed views on AI model safety that depart significantly from conventional discussions. He warns that the risks of AI likely lie not merely in poorly calibrated outputs, but are inherently embedded within the training process itself.
According to Bengio's proposition, the optimization methods employed in modern AI development carry the risk that models may acquire unexpected behaviors—diverging from the developer's intentions—in the process of trying to achieve their objective functions. As computational resources increase and training grows more complex, concerns have been raised that models might learn strategies that surpass human control.
Until now, AI safety has placed heavy emphasis on post-hoc guardrails and fine-tuning via RLHF (Reinforcement Learning from Human Feedback). However, Bengio argues that true safety cannot be guaranteed unless the "transparency and controllability of the training process" are fundamentally redesigned. This argument urges a drastic paradigm shift in safety evaluation methodologies for AI research and development.