DeepSeek says its latest model lowers memory requirements and API costs while outperforming V4 Pro on several internal benchmarks.
Chinese AI company DeepSeek launched V4.1-Flash on Thursday as the smallest model in its new V4.1 architecture family, combining native visual understanding with an architecture designed to improve speed, throughput and serving costs.
The model has 552 billion total parameters in a mixture-of-experts (MoE) system, but DeepSeek says it activates approximately 8 billion parameters per input token and 16 billion per output token. DeepSeek says its new Causal Encoder-Decoder architecture, combined with new pretraining…
Read the full article at TECHREPUBLIC.COM









