
Let's Run DeepSeek V4 Flash vs Pro - Local AI Coding, Maths & Logic TESTED π§
Keywords
Summary
153 words
Critical Evaluation
Value of the Information & Strength of the Argument
The video provides hands-on data on running a large model locally, including token speeds, memory usage, and qualitative outputs. It offers practical guidance on quantization choices. The tests are not exhaustive but give a decent overview. The creator clearly explains his methodology, tests multiple scenarios, and compares with cloud versions. However, conclusions are based on single runs and subjective observations, and the IMO question is just one instance. The argument that the Flash model is a breakthrough is somewhat anecdotally supported.
Scientific Rigor, Source Quality, Title Accuracy
The sources include links to the model and companion videos, but no official DeepSeek documentation. The title accurately describes the content. The creator is transparent about using quantized models, and points out the inferencing engine limitations, which adds credibility.
135 words
Title / Content Match
The title accurately reflects the content: the creator runs DeepSeek V4 Flash locally and compares it with Pro, testing coding, math, and logic.
Quality & Reliability
7/10
The video presents hands-on testing with clear methodology, but relies on subjective visual comparisons and does not provide reproducible benchmarks. Some claims are not externally verified.
Key Moments
Markers derived by PSI from the transcript: the creator did not define chapters.
- Introduction and overview of DeepSeek V4 editions
- Solar system test with local quantizations and comparison to cloud
- Flappy Bird test showing visual differences and token speeds
- Minecraft test evaluating generations and memory usage
- Logic and riddle tests including trolley problem and surgeon riddle
- IMO math problem solving, showing the Flash model's success
Cited Sources
- DeepSeek-V4-Flash-MLX-9bit β Quantized model version uploaded by the creator for local inference
- Inferencer App β Software used to run the model locally
- Kimi K2.6 Video β Companion video comparing another large model
- GLM 5.1 Video β Companion video comparing another large model
- Expert Controls Video β Earlier video by the creator on model control features
Concurring Sources
- DeepSeek-V4-Flash-MLX-9bit β The uploaded model itself is a direct resource for replicating the local tests.
Contribution & Novelties
The video contributes a hands-on, practical comparison of running a cutting-edge large language model locally, highlighting the trade-offs between quantization approaches (repacking vs. requantization) and their impact on memory and output quality. It also showcases the model’s coding, logic, and mathematical abilities through concrete tests. The creator openly shares an experimental quantized model for others to try.
Pour aller plus loin :
- Mixture of experts β DeepSeek models often use Mixture of Experts, a key architecture concept for scaling parameters efficiently.
- Model compression β Quantization techniques discussed in the video are a form of model compression, critical for running large models on consumer hardware.
- Transformer architecture β The foundational architecture behind DeepSeek and other modern LLMs, relevant for understanding attention and inference.
122 words
Radar Profile
The profile shows a balanced performance across quantitative information, qualitative assessment, technical depth, and overall reliability, with a slightly higher score in information quantity and lower in reliability, reflecting the informal nature of the review.