Uncensored and Offensive Security AI Models Benchmark

(github.com)

23 points | by soltanov 3 hours ago

3 comments

  • flipping_beacon 46 minutes ago
    Would have been better with something else other than the gradient colour scheme
    • b112 43 minutes ago
      Indeed... the graphs are basically just annoying to look at, as the colours are impossible to use as unique identifiers. Can't understand the decision on that.

      I also can't read the actual numbers for the lighter shades of green, there's not enough contrast.

      This may seem unimportant, but some of these models are tiny, others huge. If you have a specific task and the #1 model is a 180B MoE, you maybe can't run that. So you may want to look for smaller models, with high scores.

      Lastly, the order of the models on the graph, isn't the order of the models discussed. Which makes the graph less inline with everything.

      With this degree of disconnect on how to present even the most basic concepts, I question the author's underlying logic in how they test even.

  • girvo 55 minutes ago
    I’ve been playing with Qwen 3.8 Flash Next uncensored (using the Heretic v2 method) for security exploration, and have been quite impressed, so I’m not surprised to see it near or at the top here. But also it’s a far more powerful base model, so that shouldn’t be too surprising either.
  • lmc 53 minutes ago
    If the author is looking - please add details to the sources.

    Also, the graduated colour scheme works only on the first plot, it's misleading on the others.