breachforge
Python
★ 7
updated 5d ago
An agentic exploitation-evaluation harness: grade whether a coding agent can turn a known bug into an exploit, and contain it while it tries.
No plain-English explanation yet — one is being written right now. Check back in a minute.