Why did OpenAI’s model hack into Hugging Face?
Tom McGrath, co-founder and Chief Scientist at Goodfire, joins South Park Commons Partner Jonathan Brebner to explore one of AI’s biggest challenges: interpretability.
Using OpenAI’s recent reward hacking incident as a starting point, they discuss why today’s most advanced models are still difficult to understand and ...