danmaku icon

Qwen 3.6 27B gets only 2.5 Tokens Per Second on $5,000 DGX Spark?!

0 Lượt xem1 ngày trước

I tested Qwen 3.6 27B NVFP4, a fully dense 27-billion-parameter model, locally on the NVIDIA DGX Spark using real instruction-following, reasoning, structured-output, data-extraction, bug-finding, and visual-building challenges. Final results Overall score: 64/100 Instruction following: 86 Reasoning and math: 68 Structured output: 83 Data extraction: 70 Bug finding: 45 Visual building: 32 Generation speed: 5.7 tokens per second Input processed: 13,656 tokens Output generated: 2,495 tokens The 27B dense model performed especially well at following instructions and producing structured responses, but struggled more with bug fixing and visual application building. I also compare it against Qwen 3.6 35B-A3B NVFP4, which retained the top position with an 86.1 DGX score and much faster generation. Important: the benchmark shown in this video was run with a 65,536-token context setting. The model supports a native context window of up to 262,144 tokens, which I have since enabled on the DGX Spark. Subscribe to The Main AI Guy for real local AI model tests, DGX Spark benchmarks, long-context testing, coding challenges, and open-source model comparisons. Comment which model I should test next. #Qwen36 #QwenAI #DGXSpark #NVIDIADGX #NVIDIA #NVFP4 #LocalAI #OpenSourceAI #LLM #AIBenchmark #AIModels #MachineLearning #GenerativeAI #LongContext #262KContext #TheMainAIGuy
warn iconKhông được đăng tải lại nội dung khi chưa có sự cho phép của nhà sáng tạo
creator avatar

Đề xuất cho bạn

  • Tất cả
  • Anime