FreeToken Open-source: Running a DeepSeek-V4-Flash with a maximum of 25 Token /s
2026-09-01 15:37:47
According to CoinMeta, the FreeToken open-source project was launched by researchers from Berkeley, MIT, and others. It aims to create a hybrid expert model inference engine designed for local devices, addressing the issue of insufficient video memory in very large models. This engine stores some of the weights in system memory, allowing CPU and GPU to participate in the inference process together. Tests have shown that a desktop equipped with a RTX 5090 graphics card running a DeepSeek-V4-Flash model with 284B parameters can achieve a speed of 22–25 Token per second, while on a 8GB video memory RTX 4060 laptop, a 35B model can reach a speed of 39.3 Token per second. The project currently supports over 20 different MOE models and has been made open-source under the Apache 2.0 license.
Bullish 0
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.