Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

What happens if a model passes the government tests and then later someone fine tunes it to behave differently, without making their changes public?


Any reasonable safety testing should include finetuning and safety margin to account for others may do better finetuning.


I can fine tune significant behavior changes, there is little model developers can do to prevent this (aiui), so this effectively becomes an blanket ban


Yes, I agree it is effectively a blanket ban (above some capability) for now. I hope AI alignment research advances in the future so that it is not so.


a ban is effectively impossible without a global treaty

the current US admin as pulled out and worked against all sorts of global treaties, agreements, and negotiations; sending the president's friends instead of experts; who's going to trust us?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: