Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Open weights under the MIT license: what exactly DeepSeek is releasing

The repo and the V4.1-Flash weights are under the MIT license. DeepSeek opened its first multimodal model, V4, the same way on 31 August 2026: Flash-Vision-Exp.

TechnologyAnalysisRachel NwosuPublished: 25 September 20266 min readSources 3
Open weights under the MIT license: what exactly DeepSeek is releasing

The license is one of the strongest arguments in the model war right now. The DeepSeek-V4.1-Flash model card states it plainly: the repository and the model weights are covered by the MIT license. That is one of the most permissive open source licenses. It does not require changes to be disclosed and does not restrict commercial use.

The scale of the release is larger than it might seem. The Hugging Face repository holds 48 weight files in safetensors format, and next to them a set of tools. There is no Jinja chat template, but there is standalone reference encoding code in Python that handles multi-turn conversations, tool calls, reasoning mode, numerical reasoning effort and images woven into the text stream.

Not just the model files

The interoperability statement has value in itself. The same model card lists deployment variants for vLLM, SGLang, TensorRT and Docker Model Runner, plus FP8, BF16, GPTQ, AWG and GGUF quantizations. It also names outside inference providers, among them novita, fireworks, featherless and deepinfra, that run the model on their own hardware. The model also runs in llama.cpp, Ollama and LM Studio. The weights are not released on paper only. They can genuinely be run outside the maker's infrastructure.

DeepSeek took the earlier step in the same direction on 31 August 2026. The Chinese outlet IT之家 noted at the time that DeepSeek-V4-Flash-Vision-Exp, the first multimodal model in the V4 series, available in the API since 21 August 2026, also landed on Hugging Face under the MIT license. The company opened not only the model files and tokenizer, but also the reference prompt encoding implementation and a minimal inference implementation in PyTorch covering the vision encoder, the aligner, DFlash attention, MoE, hyper-connections and DSpark. The model accepts JPEG, PNG, GIF and WebP images; the maker labelled it an experimental version.

Openness as a commercial weapon

This is not a selfless gesture. The Italian outlet CorCom argued in an analysis dated 7 September 2026 that cheap open-weight models from China can squeeze the margins of companies building the most expensive AI systems. The authors, Filippo Pallotti and Michael Strobaek of Banque Lombard Odier, add a caveat: the advantage of Chinese models is overstated, because distillation and the cost and time of completing a task can raise the real cost of work even if the price per token looks excellent. That is an important addition to the enthusiasm around the MIT license. Free weights do not mean a free deployment.

Yet the figure of 621 396 downloads in the past month on the model card shows that the availability of the weights alone draws engineers in, regardless of the argument over margins.

Comments 0

Sources

3
  1. 01DeepSeek-V4.1-Flash — sekcja License i wdrożeniaEN
  2. 02DeepSeek-V4-Flash-Vision-Exp 模型已开源 (IT之家)ZH
  3. 03AI, la Cina sfida i big Usa sui prezzi (CorCom)IT

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.