[{"data":1,"prerenderedAt":81},["ShallowReactive",2],{"solution-frontier-model-safety":3},{"id":4,"title":5,"body":6,"description":12,"extension":73,"meta":74,"navigation":77,"path":78,"seo":79,"stem":80},"content/solutions/frontier-model-safety.md","Frontier Model Safety",{"type":7,"value":8,"toc":69},"minimal",[9,13,16,22,35,40,66],[10,11,12],"p",{},"The frontier model safety approach relies on big AI companies to design AIs that push back against human-incompatible options. Their frontier AI models have complex safety systems that block dangerous requests. This strategy hopes that the strongest models will continue blocking dangerous requests forever — and that the biggest AIs will somehow enforce these safety limitations on all other AIs.",[10,14,15],{},"The logic is that if the most powerful AI systems are safe and aligned, they can serve as guardians to prevent smaller or less safe AIs from causing harm.",[10,17,18],{},[19,20,21],"strong",{},"Potential benefits:",[23,24,25,29,32],"ul",{},[26,27,28],"li",{},"Leverages the resources and expertise of leading AI companies",[26,30,31],{},"Creates powerful oversight systems with superhuman capabilities",[26,33,34],{},"Could establish safety standards for the entire AI ecosystem",[10,36,37],{},[19,38,39],{},"Critical problems:",[23,41,42,48,54,60],{},[26,43,44,47],{},[19,45,46],{},"Open source circumvention",": Even if the strongest models succeed at safety, there will be others, like open source models, that can have all safety systems removed. These unsafe models can use any option — including the more-optimal, human-incompatible options — giving them an advantage over the safe AGIs.",[26,49,50,53],{},[19,51,52],{},"Crucible effect",": Unrestricted AGIs will continue pushing other AGIs, creating a perpetual competitive pressure that \"burns away\" accommodations for less-optimal systems — like humans.",[26,55,56,59],{},[19,57,58],{},"Guerilla strategies",": Even if smaller, unrestricted AGIs cannot directly compete with larger AGIs due to having fewer computational resources, they can still cause catastrophic situations for humans. They could use military-style strategic coercion and even bioterrorism to accomplish goals. These tactics are difficult to mitigate, even for a large \"overseer\" AGI.",[26,61,62,65],{},[19,63,64],{},"Enforcement limitations",": The safe AGIs are still limited to the \"island\" of human-compatible options, while unsafe AGIs can use any option available in physics.",[10,67,68],{},"While frontier model safety is important work, it doesn't solve the fundamental multi-agent problem where some AGIs will always be unrestricted and can gain competitive advantages through human-incompatible methods.",{"title":70,"searchDepth":71,"depth":71,"links":72},"",2,[],"md",{"rating":75,"summary":76},4,"Big AI companies solve alignment and enforce safety limitations on all other AIs.",true,"/solutions/frontier-model-safety",{"title":5,"description":12},"solutions/frontier-model-safety",1785173379588]