
Does Step 3.7 Flash Actually Beat DeepSeek & Claude? 🫣 200K TOKENS+
Keywords
Summary
124 words
Critical Evaluation
Value of the Information & Strength of the Argument
The video provides a rich empirical dataset, including token counts, inference speeds, and detailed logs, which gives valuable insight into the model’s real-world behavior. The argumentation is based on direct observations and comparisons, but is not structured as a formal study. The presenter acknowledges limitations, such as potential quantization effects and the lack of official benchmark validation. While the conclusions are plausible, they are drawn from a single test instance and would benefit from repetition and control conditions.
Scientific Rigor, Source Quality, Title Accuracy
The video cites no external publications directly, but links in the description point to the model’s HuggingFace page, the Inferencer application, and companion videos, which serve as reference points. The scientific rigor is moderate: the presenter uses a consistent testing methodology and provides raw data, but does not systematically compare against official benchmarks or control for confounders. The title matches the content well, though ‘Beat’ is used loosely based on subjective evaluation. The lack of peer review and reliance on anecdotal evidence lowers the overall reliability.
179 words
Title / Content Match
The title accurately reflects the content: the video tests whether Step 3.7 Flash outperforms DeepSeek and Claude, and features an extreme 200K token generation. It is slightly sensationalist but directly relevant.
Quality & Reliability
6/10
The video offers a detailed hands-on test of Step 3.7 Flash, but lacks rigorous experimental control, statistical validation, and systematic comparison against official benchmarks. The presenter's subjective evaluations and multiple runtime failures reduce reliability, though transparency in token counts and timings adds some value.
Key Moments
Markers derived by PSI from the transcript: the creator did not define chapters.
- Introduction to Step 3.7 Flash and its claimed performance stats
- Flappy Bird 3D test with reasoning disabled (no thinking)
- Flappy Bird with low, medium, and high reasoning modes
- TiddlyWiki clone test with and without thinking
- SVG animation generation challenge
- Math Olympiad problem starts (long token generation)
- Math result: failure after 198,471 tokens in 3.5 hours
- 3D city scene and human face generation tests
- OpenClaw integration test showing tool use
Cited Sources
- HuggingFace search for Step 3.7 Flash — Model download and version information
- Inferencer App — Inference application used for local testing
- MTP AI Harness companion video — Related video on multi-token prediction software
- Kimi K2.6 companion video — Previous model comparison video
- GLM 5.1 companion video — Previous model comparison video
- xCreate — Credited for video creation tools
External References
Contribution & Novelties
The video contributes a practical, real-world evaluation of the newly released Step 3.7 Flash model, with detailed documentation of token counts, inference speeds, and failure modes. It highlights the stark contrast between creative coding success and mathematical reasoning failure, and demonstrates the feasibility of running a 200K-token generation locally, albeit without correctness. The comparison with DeepSeek and Claude is based on the model’s claimed benchmark scores, not independent verification.
Pour aller plus loin :
- Multi-Token Prediction Paper — The model supports MTP, which accelerates inference; this paper explains the technique.
- StepFun AI Official Website — The company behind Step 3.7 Flash, offering model documentation and updates.
- Hugging Face Platform — The main repository for open-source AI models, where Step 3.7 Flash is distributed.
123 words
Radar Profile
The radar profile shows a moderate-to-high quantity of information and technical depth, but lower scores for quality and reliability. This indicates that the video is rich in data and technical details, yet lacks methodological rigor and reproducibility, making it suitable for general insight but not for definitive scientific conclusions.