Alibaba Cloud released Qwen3.8-Flash-Next, a 125B-A6B parameter model that outperforms larger competitors like DeepSeek V4 Flash despite its smaller size. Using Qwen Sparse Attention architecture, it delivers 7.6x faster prompt processing and 4.9x faster token generation while maintaining multimodal (image reading) capabilities. Priced at $0.16/0.47 per million tokens, it's competitive for local deployment use cases, though API pricing matches rivals. Relevant for Thai SMEs evaluating AI infrastructure efficiency.
← Back to all articles