
Qwen 3.8-Max Is the First Max-Tier Qwen With Open Weights, and the 27B Fits in 17GB
Quick verdict
Alibaba shipped Qwen3.8-Max and, for the first time in the Max line, said the weights are going open next week. It lands near the top of the open-weight tables, with an experimental score of 143.33 on BenchmarkList and a #12 spot out of 374 models tracked, second among open-weight models. API pricing is $2 per million input tokens and $6 per million output, with cached input at $0.25. The part local users cared about most is the smaller sibling: Qwen3.8-27B, which Daniel Han of Unsloth says should run in roughly 17GB of VRAM. If you have been waiting for a frontier-class model you can actually download, this is the closest a Max-tier release has come.
What actually shipped
Two models, announced together. Qwen3.8-Max is the big one at 2.4T total parameters. Qwen3.8-27B is the small, dense-ish model built to run on hardware people own.
- Qwen3.8-Max scores 143.33 on BenchmarkList's experimental index and ranks #12 of 374 models overall, #2 among open-weight releases. The category tables point to coding and software tasks as its strongest area.
- API pricing is $2 per million input, $6 per million output, and $0.25 per million on implicit cache hits. Alibaba framed the launch line as "better and cheaper."
- The Max weights are set to release next week. Prior Max-tier Qwen models were API-only, so this is the first time a top Qwen model is going open.
- Qwen3.8-27B is the local play. Daniel Han of Unsloth validated that it should fit in about 17GB of RAM or VRAM, which points at a quantization-aware build rather than a naive post-training quant.
- On the vision side, Qwen3.8-Max does box-conditioned detection: skalskip92 reported 60% mAP from a single box and 80% with multiple boxes on hard-to-describe concepts. Qwen-Image-3.0-Pro also reached #5 on the Text-to-Image Arena.
The blog claims are more ambitious than a benchmark card. Qwen describes 10-plus days of autonomous software development from an empty repo, a native visual feedback loop for execution and correction, and a closed-loop chip-design pipeline running Iverilog, Yosys, and OpenROAD over 500-plus turns. In that run it reports cutting a crypto accelerator from 8,298 gates down to 678 while still hitting timing closure. Some readers on Reddit called the framing jargon-heavy, and it is worth treating the agentic claims as vendor marketing until someone reproduces them.
The ecosystem moved fast, which is usually the real signal with a Qwen drop. Nous Research and Cline both wired the model into their agent stacks within a day, and the community started lining Qwen3.8-Max up against Kimi K3 and DeepSeek-V4-Flash on the open-weight leaderboards almost immediately. One caveat: the 27B benchmarks are not out yet, so its quality is still an open question even though its memory footprint is known.
Why it matters
The Max-tier weights going open is the headline. Until now the strongest Qwen model was something you rented through an API. Opening it changes the math for anyone who wants a frontier-class model without a per-token bill or a vendor dependency. That is the same pressure DeepSeek has been applying, and we walked through the last round of it in DeepSeek V4-Flash shipping open weights the same day. Two of the strongest open labs are now racing to give away better models, which puts a floor under what anyone can charge for this tier.
The 27B is the part most people will actually run. A 17GB footprint means it fits on a single consumer GPU, which is the difference between a model you read about and a model you use. Reddit was split on that number, since 17GB narrowly excludes common 16GB cards, but for anyone on a 24GB GPU it opens a genuinely capable local model. That matters most for high-volume or background work where you do not want to pay per call. The practical move is rarely "run everything locally," though. It is routing the cheap tier to the tasks that fit it and saving the paid frontier for the rest, which is the case we make in cutting AI coding agent costs with model routing.
The wider pattern is that open-weight models keep closing the gap on the paid frontier faster than paid plans get cheaper. A Max-tier model going open, plus a 27B that runs on one card, would have been two separate headlines a year ago. Now they ship on a quiet Tuesday. We track where that leaves the open ecosystem in open-source AI models in 2026 and how the Chinese labs are diverging in Chinese open models climbing OpenRouter. If you would rather not commit to one lab, running several models behind one subscription is covered in the best app for running multiple AI models.
Video: Qwen 3.8-Max and the open-weight question
This walks through the Qwen3.8-Max launch, the benchmark claims, and whether it holds up as an open model against Kimi K3 and DeepSeek.
FAQ
Are the Qwen3.8-Max weights actually open?
Alibaba said the Max weights release next week, which would make it the first Max-tier Qwen available to download rather than only rent. Until they land, the Max model is API-only at $2 per million input and $6 per million output. The 27B is the model built for local use.
What hardware do I need for Qwen3.8-27B?
Daniel Han of Unsloth put it at roughly 17GB of RAM or VRAM, which suggests a quantization-aware build. That fits a single 24GB consumer GPU comfortably but narrowly misses 16GB cards. The 27B benchmarks are not out yet, so test it on your own tasks before trusting the size alone.
How does it compare to DeepSeek V4-Flash and Kimi K3?
The community is lining all three up on the open-weight leaderboards, with Qwen3.8-Max at 2.4T total parameters against Kimi K3's 2.8T and DeepSeek-V4-Flash's much smaller 284B. On raw score they are close, though a sub-300B model holding its own is part of why DeepSeek keeps drawing attention. For that side of the race, see our writeup of DeepSeek V4-Flash.
Should I switch my coding agent to it?
Worth testing rather than trusting the launch charts, especially since the agentic claims in the blog are unverified. It is already in Cline and Nous stacks, so the cost of trying it is low. The tier that fits depends on your work, which is the point of comparing agents in the best AI coding agents guide.
Sources
- @Alibaba_Qwen - Qwen3.8-Max launch, "better and cheaper"
- @Alibaba_Qwen - Qwen3.8-27B announced alongside Max
- @skalskip92 - box-conditioned detection, 60% and 80% mAP
- @arena - Qwen-Image-3.0-Pro reaches #5 in the Text-to-Image Arena
- r/LocalLLaMA - Daniel Han validates Qwen3.8-27B at ~17GB
- r/LocalLLaMA - BenchmarkList score, rank, and pricing
- Qwen - official Qwen3.8 blog and agentic claims
Further reading
Try all the models mentioned in this article
Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.
Start free on Admix