Six Weeks of Releases, in One Table
| Model | Developer | Released | Why it matters |
|---|---|---|---|
| DeepSeek V4 Flash 0731 | DeepSeek | Jul 31, 2026 | Huge benchmark jump over prior Flash; best price/performance of the summer |
| Kimi K3 | Moonshot AI | Jul 16 (weights Jul 27) | Flagship continuation of the Kimi K2 lineage — strong open coding/reasoning |
| Gemma 3 (27B) | Google DeepMind | Jul 2026 | 27B-class open model; local-friendly at Q4 on 24GB |
| Mistral Medium 3 | Mistral AI | Jul 2026 | EU champion's mid-tier refresh |
| GLM-5.2 | Z.ai / Zhipu | Jun 13, 2026 | Open leader in agentic capability; beats DeepSeek V4 Pro on shared SWE benchmarks |
| Qwen Image 3.0 / 3.0 Pro | Alibaba (Qwen) | Aug 2026 | Open text-to-image family — first of August's releases |
| Muse Spark 1.2 | Tencent | Aug 2026 | Image-gen update; the visual-model race is on |
The Headline: DeepSeek V4 Flash 0731
DeepSeek's mid-summer update turned the value conversation upside down. The 0731 build of V4 Flash claims ~79% SWE-bench Verified and 91.6% LiveCodeBench — numbers that would have been frontier-class a year ago — at a price point under everything comparable. The positioning is explicit: frontier-ish quality, commodity pricing. For high-volume coding workloads and agent fleets, this is the model that makes API cost a rounding error.
Caveat that matters: it's an API-first MoE. Self-hosting V4 Flash needs serious VRAM, so local users should treat this as the price argument for open AI while the sub-12B line remains the hardware story. Our GLM vs DeepSeek deep dive covers the family matchup.
Kimi K3: Moonshot's Next Salvo
Kimi K3 (announced July 16, weights July 27) continues Moonshot's K-series — the lineage that made Kimi K2 a favorite among open-weight coding models. The K3 launch keeps the pattern: strong reasoning, strong code, open weights, aggressive pricing. It's part of the four-way open-frontier race — GLM, DeepSeek, Kimi, Qwen — where each release leapfrogs the last within weeks. Open beating closed is no longer a slogan; it's a release cadence.
What of This Is Local-Runnable?
| Model | Local reality |
|---|---|
| Gemma 3 27B | ✅ Q4 fits a 24GB GPU — the practical self-host pick of the batch |
| Mistral Medium 3 | 🟡 48GB-class machine, or API |
| DeepSeek V4 Flash / Kimi K3 / GLM-5.2 | 🔴 MoE giants — API or enterprise hardware only |
| Qwen Image 3.0 | 🟡 Weights open; VRAM-heavy for image generation |
The honest local takeaway: the most interesting releases of the summer are not the ones you can run. If you're self-hosting, the practical upgrades remain the RAM-ranked GGUF picks and sub-12B coding models — while the giants set your expectations for quality.
The Three Trends This Roundup Confirms
- Open is catching up, permanently. Open-weight models now match closed ones on most benchmarks — see the 2026 edition — and this summer's releases only widen the gap.
- MoE is the frontier architecture. Every giant (V4, K3, GLM-5.x) is a mixture-of-experts with a small active-parameter count — efficiency is the arms race now, not just capability.
- Speed of iteration beats individual launches. Flash updates (0731), point releases (5.2), and refresh cycles matter more than any single debut. Bookmark a timeline tracker and check monthly.
Frequently Asked Questions (FAQ)
What were the biggest AI model releases in July–August 2026?
The standouts: DeepSeek V4 Flash 0731 (July 31 — huge benchmark jumps at low cost), Kimi K3 (Moonshot, weights July 27), GLM-5.2 (June 13, the agentic open leader), Gemma 3 27B (July), Mistral Medium 3 (July), and Qwen Image 3.0 Pro (August, image generation).
Is DeepSeek V4 Flash 0731 worth using?
Yes for budget-maximized workloads — the 0731 update delivered large jumps (roughly 79% SWE-bench Verified, 91.6% LiveCodeBench) at a dirt-cheap price, positioning it as the best value model of the summer. It's an API model first; self-hosting the MoE requires serious VRAM.
What is Kimi K3?
Moonshot AI's newest flagship, announced July 16, 2026 with weights released July 27. It continues the Kimi K2 lineage of strong open-weight coding and reasoning models and is part of the open-frontier wave alongside GLM and DeepSeek.
Which August 2026 releases can I run locally?
Gemma 3 27B fits a 24GB GPU at Q4; Mistral Medium 3 runs on a 48GB-class machine or via API. The MoE giants (DeepSeek V4 Flash, Kimi K3, GLM-5.2) are impractical to self-host for most users. For laptops, the recent practical option remains the Qwen3.5 and Gemma 4 sub-12B line.
What is Qwen Image 3.0?
Alibaba's new open image-generation model family, released in August 2026 (standard and Pro variants). It extends the Qwen family into competitive text-to-image, continuing Alibaba's pattern of releasing strong open weights.
What should I watch next?
The open-vs-closed race: open weights (GLM, DeepSeek, Kimi, Qwen) are matching closed models on most benchmarks while costing a fraction. Watch for the next Qwen 3.x, the next Kimi K-series, and whether any vendor ships a self-hostable 100B+ that fits 48GB.
Sources & Further Reading
- GLM vs DeepSeek: Which Open Model Family in 2026?
- Open Models That Beat Closed Ones: 2026 Edition
- Top 10 Open Source Models
- Top 10 AI Models Under 12B Parameters
- New AI Model Releases — August 2026 timeline (LLM Gateway)
- 9 open-source launches in 12 days (Tech Insider, Jul 2026)
- DeepSeek V4 Flash 0731 tested (Jul 31, 2026)