bitcoin
bitcoin

$77625.828729 USD

0.62%

ethereum
ethereum

$2517.417853 USD

0.13%

tether
tether

$0.999544 USD

-0.01%

bnb
bnb

$723.660100 USD

0.16%

xrp
xrp

$1.386011 USD

1.82%

usd-coin
usd-coin

$0.999860 USD

0.01%

solana
solana

$101.519022 USD

0.20%

tron
tron

$0.339380 USD

-0.12%

hyperliquid
hyperliquid

$79.865341 USD

1.22%

zcash
zcash

$1136.272189 USD

-0.43%

dogecoin
dogecoin

$0.084270 USD

-0.24%

monero
monero

$505.727561 USD

-4.73%

chainlink
chainlink

$11.414119 USD

-0.35%

unus-sed-leo
unus-sed-leo

$8.961937 USD

-1.04%

cardano
cardano

$0.209002 USD

0.75%

Cryptocurrency News Video

DeepSeek-V4.1-Flash is online: supports 1M token context. global KV cache 890 bytes per token

Sep 11, 2026 at 02:20 am easyvibecoding

DeepSeek-V4.1-Flash is online: supports 1M token context, global KV cache 890 bytes per token. Release Highlights DeepSeek announced DeepSeek-V4.1-Flash on September 10, 2026, which is positioned as the smallest model in the new architecture family, supports native image and text processing, and claims to have improved capabilities, inference speed, throughput, and scalability to larger models. The model has been launched on the DeepSeek API. Set the model to deepseek-flash to use native multi-modal support. Architecture and Efficiency DeepSeek-V4.1-Flash is a multi-modal MoE with 552B backbone parameters and supports up to 1M token context. Its Causal Encoder-Decoder (CED) architecture consists of 20 layers of causal encoder and 20 layers of decoder. Each input prefill token enables 8B parameters, and the output decode enables 16B parameters. Each layer of MoE of the model uses 1 shared expert and 384 routed experts, and each token enables 6 routed experts. The model is trained from scratch with multi-modal data of 45T tokens, sparse attention is first trained with 64K sequence length, and then the context is expanded to 1M tokens at 34T tokens. Compressed Sparse Attention 2 (CSA2), Hierarchical Sparse Indexer and FP4 main KV caching compress the global KV cache to 890 bytes per token, which is about 1/4 of DeepSeek-V4-Flash. SWA Bounded Replay only rebuilds the latest n_win token, reducing the persistent KV cache to about 1/8 of DeepSeek-V4-Flash; the announcement summarizes that HBM requirements are 1/4 of the previous generation and SSD storage requirements are 1/8. The inference control and agent workload model provides an integer reasoning effort from 1 to 100 to continuously adjust the trade-off between reasoning cost and accuracy. Officials link KV cache compression to Agent costs, pointing out that cache-hit costs often account for a large proportion of Agent costs, so shrinking the cache can reduce service costs for long-context workloads. The model also integrates components such as Engram conditional memory, DSpark speculative decoding and Single-Pass mHC. Post-training follows the SFT → RL → on-policy distillation process. The main changes focus on automatically synthesizing Agent tasks, environments and rollout data pipelines. The evaluation results and conditions are officially compared with base models in DeepSeek's internal framework and under the same evaluation settings. Score differences within 0.3 are considered equivalent. DeepSeek-V4.1-Flash-Base scored 74.1 in MMLU-Pro, 60.6 in BigCodeBench Pass@1, 79.4 in HumanEval Pass@1, and 93.0 in GSM8K; comparisons of these results with DeepSeek-V4-Flash-Base and DeepSeek-V4-Pro-Base are all internal framework evaluations. In the Agent scaffold test, all settings used Linux containers, temperature=1.0, topp=0.95, 1M token context limit, maxsteps=500 and Max reasoning effort; DeepSWE v1.1 used N=8 per task, Terminal-Bench 2.1 used N=3, and the latter did not provide a network. DeepSeek-V4.1-Flash combined with DSH Minimal scored 72.6 in DeepSWE v1.1 and 90.6 in Terminal-Bench 2.1 Pass@1; mini-SWE-agent scored 74.2 and 90.3 respectively. DeepSeek-V4.1-Flash leads other models in DeepSWE v1.1 (74.2), CyberGym (88.1) and Automation-Bench (54.8) scores, but lags behind Opus5 and GPT5.6-Sol in Terminal-Bench 3.0 (30.0). API, Compatibility and Deployment V4-Flash and V4-Flash-Vision-Exp have been retired; deepseek-v4-flash and deepseek-v4-flash-vision-exp are currently only temporarily routed to V4.1-Flash. This compatibility route is not a permanent arrangement. The read source snippet shows that API is priced at peak/off-peak prices, and the off-peak price is 50% of the peak price; the full price has not yet been confirmed. DeepSeek also stated that it will work with the open source community to promote V4.1-Flash inference support and more deployment options; currently the encoding folder and deepseek-recipe provide prompt encoding and API format conversion tools, but model inference, tool execution and HTTP transport are still the responsibility of the caller. Original text: https://easyvibecoding.app/curated/3303-deepseek-releases-deepseek-v4-1-flash-1m-token-context-kv
Video source:Youtube

Disclaimer:info@kdj.com

The information provided is not trading advice. kdj.com does not assume any responsibility for any investments made based on the information provided in this article. Cryptocurrencies are highly volatile and it is highly recommended that you invest with caution after thorough research!

If you believe that the content used on this website infringes your copyright, please contact us immediately (info@kdj.com) and we will delete it promptly.

Other videos published on Sep 15, 2026