I tested Qwen 3.6 27B NVFP4, a fully dense 27-billion-parameter model, locally on the NVIDIA DGX Spark using real instruction-following, reasoning, structured-output, data-extraction, bug-finding, and visual-building challenges.
Final results
Overall score: 64/100
Instruction following: 86
Reasoning and math: 68
Structured output: 83
Data extraction: 70
Bug finding: 45
Visual building: 32
Generation speed: 5.7 tokens per second
Input processed: 13,656 tokens
Output generated: 2,495 tokens
The 27B dense model performed especially well at following instructions and producing structured responses, but struggled more with bug fixing and visual application building. I also compare it against Qwen 3.6 35B-A3B NVFP4, which retained the top position with an 86.1 DGX score and much faster generation.
Important: the benchmark shown in this video was run with a 65,536-token context setting. The model supports a native context window of up to 262,144 tokens, which I have since enabled on the DGX Spark.
Subscribe to The Main AI Guy for real local AI model tests, DGX Spark benchmarks, long-context testing, coding challenges, and open-source model comparisons.
Comment which model I should test next.
#Qwen36 #QwenAI #DGXSpark #NVIDIADGX #NVIDIA #NVFP4 #LocalAI #OpenSourceAI #LLM #AIBenchmark #AIModels #MachineLearning #GenerativeAI #LongContext #262KContext #TheMainAIGuy