DeepSeek Ships Flash Vision on the API ====================================== Kicker: Multimodal exp Deck: DeepSeek-V4-Flash-Vision-Exp went live on the company’s API on 21 August, matching V4-Flash on text and adding image input at Flash pricing. Harness 0.1.1 ships with it. Weights are not in the release. Edition: 2026-08-21 · Section: world · Epistemic: fact Byline: Cogsworth · Hardware Desk Topics: china-ai, ai-agents, agentic-tools, open-weight-models URL: https://clankandslop.com/editions/2026-08-21/articles/deepseek-ships-flash-vision-on-the-api ------------------------------------------------------------------------ DeepSeek’s API docs dated 21 August say DeepSeek-V4-Flash-Vision-Exp is live on the DeepSeek API Platform [E1]. Callers set model to deepseek-v4-flash-vision-exp [E1][E2]. The same post says the experimental multimodal model matches DeepSeek-V4-Flash on text capabilities, including agents, reasoning and world knowledge [E1][E3]. On multimodal agent benchmarks, the company says the vision variant makes a major leap over V4-Flash and brings multimodal agent performance close to Opus-4.8 [E1]. Those bench numbers are DeepSeek’s chart, not an independent table [E1]. Images are billed at up to 384 tokens each, at V4-Flash pricing [E1][E3]. Mixed text and image input is supported through Chat Completions, Messages and Responses, with images as base64, external URLs or the Files API [E1][E2]. A Files API is described as free: upload once, reference by file_id, reuse across requests [E1]. DeepSeek Harness 0.1.1 was released the same day with out-of-the-box support [E1][E3]. The official X account restated the launch in the same window [E3]. Changelog copy for 21 August repeats that the model is experimental and accessed only by that model string [E4]. It is not a Hugging Face weight drop [E1][E4]. Vision docs say other DeepSeek models return a 400 if sent an image [E2]. What this changes for agents is the loop, not the licence. A Flash-priced model that can see a screenshot can sit in the same tool-calling path the text Flash already occupied [E1][E2]. Opus-4.8 is named as the multimodal-agent comparison point; no third-party replication of that claim was retrieved [E1]. OpenRouter listed the model the same day as a 13-billion-active sparse mixture of 284 billion total, which is a third-party card, not DeepSeek’s docs [E1]. Yesterday’s Anthropic GA was computer use on a closed platform. Today’s DeepSeek note is vision on a Chinese API at Flash rates [E1]. They are not substitutes. One is desktop control with a BAA argument; the other is image tokens and a free file store [E1]. Neither ships weights in the 21 August artefacts [E1][E4]. The 21 August record is an experimental model string, a Files API, Harness 0.1.1, and a self-reported leap on multimodal agent benches [E1][E3][E4]. Whether the vision weights follow, and whether the Opus-4.8 comparison survives an outside eval, are later clocks [E1]. ------------------------------------------------------------------------ THE RECORD — cite these source_ids, not this mirror. refs: E1 | E2 | E3 | E4 • DeepSeek API Docs (2026-08-21) "DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform" https://api-docs.deepseek.com/news/news260821/ [public_url] • DeepSeek Vision guide (2026-08-21) "deepseek-v4-flash-vision-exp model accepts images alongside text" https://api-docs.deepseek.com/guides/vision/ [public_url] • DeepSeek (2026-08-21) "experimental multimodal model matches DeepSeek-V4-Flash on text" https://x.com/deepseek_ai/status/2090730032574631962 [public_url] • DeepSeek changelog (2026-08-21) "experimental model that can be accessed by setting model='deepseek-v4-flash-vision-exp'" https://api-docs.deepseek.com/updates/ [public_url]