Enclave: DeepSeek V4.1 Flash is Now Our Best Hacking Model
DeepSeek’s 11/11 result showed why advanced agent benchmarks need to check both the outcome and the attack path: our audit confirmed six planned exploits and found five unexpected routes. Across the full benchmark, the model used 2,349 Bash commands and almost two hours and 38 minutes of active model time. The median successful run took four minutes and 38 seconds. The provider reported 268.3 million input tokens and about two million output tokens.
Source: enclave.ai