New Open Source Qwen 397B BEATS GLM 5.1 & Claude? 🤯 | Nex N2 Pro TESTED

New Open Source Qwen 397B BEATS GLM 5.1 & Claude? 🤯 | Nex N2 Pro TESTED

🎙 xCreate 👥 26K 📅 June 7, 2026 ⏱ 13 min 👁 10K 📄 news review 🧭 2026-09-09
Available in: English (current) Français

Keywords

Qwen 397BNex N2 ProGLM 5.1BenchmarkReasoning Loops

Summary

In this video, the creator tests the newly released open-source model Nex N2 Pro, which is based on Qwen 397B and claims to outperform GLM 5.1 and Claude Opus. The video begins by showcasing benchmark scores that support these claims, but then conducts a series of hands-on tests including 3D voxel image generation, animation, math problem solving, photorealistic face rendering, Minecraft clones, and Flappy Birds. The results are mixed: while Nex N2 Pro demonstrates impressive capabilities in generating a human face and understanding context, it suffers from frequent reasoning loops when thinking mode is enabled, and its coding outputs are often slower and less refined than expected. The video compares it directly against the original Qwen and highlights the improvements, but also notes that it does not consistently beat GLM 5.1 in practical tests. Ultimately, the creator acknowledges the potential of Nex N2 Pro but remains skeptical about its benchmark claims. The video also mentions the broader context of Qwen moving away from open source, making Nex-AGI’s efforts noteworthy.

169 words

Critical Evaluation

Value of the Information & Strength of the Argument

VALUE OF INFORMATION & SOLIDITY OF ARGUMENTATION. The video provides a practical, hands-on evaluation of a new open-source model, which is valuable given the model’s recent release. The creator runs multiple diverse tests that illustrate the model’s strengths and weaknesses, offering viewers a real-world perspective beyond static benchmarks. However, the argumentation is largely anecdotal; the creator relies on his own observations and does not perform rigorous statistical analysis or controlled experiments. The claim that the model beats GLM 5.1 is based primarily on the model’s self-reported benchmarks, which are not independently verified. While the video attempts to test this claim with coding tasks, the results are not conclusive. The reasoning loops and occasional failures suggest the model is not yet on par with leading competitors, despite some impressive outputs. Thus, the value of information is moderate, but the argumentation lacks scientific rigor.

Scientific Rigor, Source Quality, Title Accuracy

SCIENTIFIC RIGOR, QUALITY OF SOURCES, ADEQUATION OF TITLE. The video does not provide any external citations or references beyond the model’s own documentation and benchmarks. The sources used are the Hugging Face page for the model and the inferencer.com app for testing, but no independent studies or expert analyses are cited. The quality of sources is therefore limited and relies on the creators’ claims. The title poses a question about the model beating GLM 5.1 and Claude, which is somewhat misleading because the content does not demonstrate a definitive win; rather, it shows potential in some areas and clear weaknesses in others. The question mark may mitigate this, but the title still overstates the comparison. The video could benefit from more rigorous benchmarking and multiple runs to ensure reliability. Overall, the rigorousness is low to moderate.

293 words

Title / Content Match

The title asks if the model beats GLM and Claude, and while the tests show some promising results, they are inconsistent, making the title somewhat overstated but with a question mark it is not false.

Quality & Reliability

6/10

The video presents hands-on tests and benchmark scores, but relies on the model's self-reported benchmarks and anecdotal evaluations.

Key Moments

Cited Sources

External References

Contribution & Novelties

APPORT NOUVEAUTES: The video offers a first-hand, practical evaluation of the Nex N2 Pro model, which is only recently released. It demonstrates that while the model excels in some tasks like generating a photorealistic human face, it suffers from severe reasoning loops when thinking is enabled, potentially limiting its practical use. This is an important observation for potential users. Additionally, the video highlights the broader ecosystem issue of Qwen moving away from open source and the role of Nex-AGI in continuing development.

Pour aller plus loin :

124 words

Radar Profile

The radar profile shows a relatively balanced performance with moderate to high scores on information quantity and technical level, but lower scores on information quality and reliability. This suggests the video provides a good amount of technical detail but may lack depth in analysis and depends on anecdotal evidence.

Reliability 6/10