DFlash 2: Keep Drafting Parallel

(inco.ai)

54 points | by mike-the-brain 2 hours ago

5 comments

  • ilc 14 minutes ago
    Watch the video carefully. DFlash2's tool call fails on python syntax.

    Usually models in this class nail things like that 1 shot, which the other side did.

    I don't know the cause. It may be nothing. But I'd like to see the model doing something where its path is a bit more constrained, to help out rule out such oddities.

  • hypfer 2 hours ago
    Amazing tech

    > An agent writes in an afternoon what a chatbot writes in a month

    But can you just.. not.

    Your tech is so good, it speaks for itself. Don't ruin that.

  • adefa 2 hours ago
    I'm getting around 27 tokens per second decode using vLLM + Qwen 3.8 27b nvfp4 + DFlash 2 on the DGX Spark.
  • sarjann 1 hour ago
    Great news, has made low memory bandwidth model usage so much nicer.
  • verdverm 2 hours ago