Local model Qwen3.8 27B is tested on a self-made cybersecurity benchmark with 19 tasks across 6 models, finding its cyber capability is high and it rarely refuses actions in an isolated docker box.
Read the original at www.reddit.com→Hey local AI community, I've been working on this for a while and finally feel ok sharing it. It's a cyber benchmark where the model gets a shell in an isolated docker box and has to find the exact flag. Pwn, web,...
Original headline: "I made my own cybersecurity benchmark and ran Qwen3.8 27B, here's how a local model actually does at hacking"
Coverage timeline
- Oct 2, 12:15 UTC r/LocalLLaMA lead source I made my own cybersecurity benchmark and ran Qwen3.8 27B, here's how a local model actually does at hacking