Does Step 3.7 Flash Actually Beat DeepSeek & Claude? 🫣 200K TOKENS+

Does Step 3.7 Flash Actually Beat DeepSeek & Claude? 🫣 200K TOKENS+

🎙 xCreate 👥 26K 📅 May 31, 2026 ⏱ 17 min 👁 7K 📄 expert opinion 🧭 2026-09-09
Available in: English (current) Français

Keywords

Step 3.7 FlashDeepSeekClaudelocal inferencetoken generation

Summary

The presenter tests the Step 3.7 Flash model from StepFun AI on a Mac Studio with an M3 Ultra 512GB, using the Inferencer app. The model is evaluated across multiple challenges: Flappy Bird 3D generation, a TiddlyWiki clone, a car-wash scenario, SVG animation, a math Olympiad problem, a 3D city scene, and human-face generation. The model shows strong creative coding abilities, producing visually impressive results in low-thinking modes, but fails on the math problem even after generating 198,471 tokens over 3.5 hours. The video also compares local vs cloud results, noting the cloud version is faster but less detailed. Integration with OpenClaw demonstrates basic tool-use capability. Overall, the model is competitive in some domains but unreliable in complex reasoning tasks, with occasional runtime errors.

124 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides a rich empirical dataset, including token counts, inference speeds, and detailed logs, which gives valuable insight into the model’s real-world behavior. The argumentation is based on direct observations and comparisons, but is not structured as a formal study. The presenter acknowledges limitations, such as potential quantization effects and the lack of official benchmark validation. While the conclusions are plausible, they are drawn from a single test instance and would benefit from repetition and control conditions.

Scientific Rigor, Source Quality, Title Accuracy

The video cites no external publications directly, but links in the description point to the model’s HuggingFace page, the Inferencer application, and companion videos, which serve as reference points. The scientific rigor is moderate: the presenter uses a consistent testing methodology and provides raw data, but does not systematically compare against official benchmarks or control for confounders. The title matches the content well, though ‘Beat’ is used loosely based on subjective evaluation. The lack of peer review and reliance on anecdotal evidence lowers the overall reliability.

179 words

Title / Content Match

The title accurately reflects the content: the video tests whether Step 3.7 Flash outperforms DeepSeek and Claude, and features an extreme 200K token generation. It is slightly sensationalist but directly relevant.

Quality & Reliability

6/10

The video offers a detailed hands-on test of Step 3.7 Flash, but lacks rigorous experimental control, statistical validation, and systematic comparison against official benchmarks. The presenter's subjective evaluations and multiple runtime failures reduce reliability, though transparency in token counts and timings adds some value.

Key Moments

Cited Sources

External References

Contribution & Novelties

The video contributes a practical, real-world evaluation of the newly released Step 3.7 Flash model, with detailed documentation of token counts, inference speeds, and failure modes. It highlights the stark contrast between creative coding success and mathematical reasoning failure, and demonstrates the feasibility of running a 200K-token generation locally, albeit without correctness. The comparison with DeepSeek and Claude is based on the model’s claimed benchmark scores, not independent verification.

Pour aller plus loin :

123 words

Radar Profile

The radar profile shows a moderate-to-high quantity of information and technical depth, but lower scores for quality and reliability. This indicates that the video is rich in data and technical details, yet lacks methodological rigor and reproducibility, making it suitable for general insight but not for definitive scientific conclusions.

Reliability 5/10