The asymmetry of access to Fable-class models for attackers and defenders creates a dangerous gradient.
- The asymmetry of access to Fable-class models for attackers and defenders creates a dangerous gradient.
- Hugging Face at the beginning of the attack: "The platform's security team were initially stymied in their incident response (IR) by unnamed US LLM frontier model guardrails 'which cannot distinguish an incident responder from an attacker,' they said."
- A cynical take from a HackerNews comment: "The product strategy of 'consumer-grade' AI making deliberately insecure software, and then selling you limited access to the model that can fix it (if they think you deserve to pay them) is just diabolical."
- I think that take is too cynical, but it is the equilibrium the incentives will pull towards.
- This recent paper shows that code produced by every major software model contained security vulnerabilities.