In this video, I test Poolside’s Laguna S 2.1 locally on a single NVIDIA DGX Spark.
Laguna S 2.1 is a 118B-parameter Mixture-of-Experts model with roughly 8B parameters activated per token. I run the native NVFP4 checkpoint using vLLM with DFlash speculative decoding, then put the model through practical BenchLocal tests.
I cover:
Model download and setup
Loading Laguna S 2.1 with vLLM
DGX Spark compatibility
Startup time and performance
Instruction-following tests
Coding and reasoning capability
Local AI advantages and limitations
Whether this 118B model is actually practical on one machine
This is a fully local test—no cloud inference and no external API generating the model responses.
System used:
NVIDIA DGX Spark
128 GB unified memory
Laguna S 2.1 NVFP4
vLLM 0.25.1
DFlash speculative decoding
BenchLocal testing interface
Poolside publishes Laguna S 2.1 as an open-weight model under the OpenMDW-1.1 license.
Subscribe for more real-world tests of large open models running locally on the DGX Spark.
#LagunaS21 #PoolsideAI #DGXSpark #LocalAI #OpenSourceAI #OpenWeights #vLLM #AIModels #CodingAI #NVIDIA #LLM #ArtificialIntelligence #BenchLocal #TechYouTube