Can a 35B model really beat a 120B model in a practical coding test?
I ran OpenAI GPT-OSS 120B and Ornith 1.0 35B locally on my NVIDIA DGX Spark and tested them on real browser-app development—not a traditional benchmark.
Both models were challenged to build and test:
🎮 A complete browser game called Meteor Escape
🧮 A functional calculator with decimals, memory controls, history, keyboard support, and automated testing
GPT-OSS 120B eventually produced a working game, but it needed additional intervention after initially failing to create the required files and incorrectly reporting that a headless browser was unavailable. Its calculator also produced a major decimal error: 100.2 = NaN.
Despite having far fewer parameters, Ornith 1.0 35B delivered the stronger overall result in this test.
This does not prove that Ornith is better at everything—but it does show that a larger parameter count does not automatically mean better real-world performance.
Watch the full comparison to see the prompts, coding process, automated tests, generated applications, and unexpected final result.
Subscribe for more local AI model testing, coding-agent comparisons, and DGX Spark experiments.
Hashtags
#GPTOSS #OrnithAI #LocalAI #LocalLLM #DGXSpark #OpenSourceAI #AICoding #CodingAgents #LLMComparison #NVIDIAAI #OpenAI #ArtificialIntelligence
#AI #ArtificialIntelligence #AITools #ChatGPT #AIBeginners