Let's Run DeepSeek V4 Flash vs Pro - Local AI Coding, Maths & Logic TESTED 🧐

Let's Run DeepSeek V4 Flash vs Pro - Local AI Coding, Maths & Logic TESTED 🧐

πŸŽ™ xCreate πŸ‘₯ 26K πŸ“… April 27, 2026 ⏱ 22 min πŸ‘ 39K πŸ“„ tutorial 🧭 2026-09-09
Available in: English (current) FranΓ§ais

Keywords

DeepSeek V4local inferencequantizationcoding testmathematics

Summary

The video presents a practical evaluation of DeepSeek V4 Flash, a locally runnable quantization of the V4 model, and compares it with the cloud-based Pro (expert) version. The creator demonstrates running the model on a Mac Studio with 512GB RAM, using two quantizations: a Q9 and a repacked 4.4-bit version. He tests the model on generating a 3D solar system, a Flappy Bird clone, and a Minecraft-style voxel world, assessing visual quality, runtime errors, and memory usage. The local model performs comparably to the cloud versions, with the repacked quantization being more memory-efficient. He then tests the model’s logic with riddles and its mathematical abilities with an IMO problem, noting that the Flash edition solves the problem correctly with lower memory footprint than competitors. The creator also highlights architecture innovations like hybrid attention and ‘manifold constrained hyperconnections’. Overall, the video offers valuable insights for running large models locally, despite some subjective visual assessments.

153 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides hands-on data on running a large model locally, including token speeds, memory usage, and qualitative outputs. It offers practical guidance on quantization choices. The tests are not exhaustive but give a decent overview. The creator clearly explains his methodology, tests multiple scenarios, and compares with cloud versions. However, conclusions are based on single runs and subjective observations, and the IMO question is just one instance. The argument that the Flash model is a breakthrough is somewhat anecdotally supported.

Scientific Rigor, Source Quality, Title Accuracy

The sources include links to the model and companion videos, but no official DeepSeek documentation. The title accurately describes the content. The creator is transparent about using quantized models, and points out the inferencing engine limitations, which adds credibility.

135 words

Title / Content Match

The title accurately reflects the content: the creator runs DeepSeek V4 Flash locally and compares it with Pro, testing coding, math, and logic.

Quality & Reliability

7/10

The video presents hands-on testing with clear methodology, but relies on subjective visual comparisons and does not provide reproducible benchmarks. Some claims are not externally verified.

Key Moments

Cited Sources

  • DeepSeek-V4-Flash-MLX-9bit β€” Quantized model version uploaded by the creator for local inference
  • Inferencer App β€” Software used to run the model locally
  • Kimi K2.6 Video β€” Companion video comparing another large model
  • GLM 5.1 Video β€” Companion video comparing another large model
  • Expert Controls Video β€” Earlier video by the creator on model control features

Concurring Sources

  • DeepSeek-V4-Flash-MLX-9bit β€” The uploaded model itself is a direct resource for replicating the local tests.

Contribution & Novelties

The video contributes a hands-on, practical comparison of running a cutting-edge large language model locally, highlighting the trade-offs between quantization approaches (repacking vs. requantization) and their impact on memory and output quality. It also showcases the model’s coding, logic, and mathematical abilities through concrete tests. The creator openly shares an experimental quantized model for others to try.

Pour aller plus loin :

  • Mixture of experts β€” DeepSeek models often use Mixture of Experts, a key architecture concept for scaling parameters efficiently.
  • Model compression β€” Quantization techniques discussed in the video are a form of model compression, critical for running large models on consumer hardware.
  • Transformer architecture β€” The foundational architecture behind DeepSeek and other modern LLMs, relevant for understanding attention and inference.

122 words

Radar Profile

The profile shows a balanced performance across quantitative information, qualitative assessment, technical depth, and overall reliability, with a slightly higher score in information quantity and lower in reliability, reflecting the informal nature of the review.

Reliability 6/10