White-hat hacking is defense to black-hat hacking; models’ refusals complicate how companies patch security, says Dario
Read the original at old.reddit.com→You patch security holes by intentionally finding them. If the models refuse to do it, how can companies protect themselves against rogue AIs, whether they are Chinese or OpenAI/Anthropic themselves? The Hugging Face...
Original headline: "White-hat hacking IS the defense to black-hat hacking. The techniques are the same. How does Dario expect companies to do it if their models refuse?"
Coverage timeline
- Jul 28, 18:31 UTC r/LocalLLaMA lead source White-hat hacking IS the defense to black-hat hacking. The techniques are the same. How does Dario expect companies to do it if their models refuse?