The Chinese AI lab’s near-frontier model targets developers and enterprises seeking strong performance at lower costs than Fable 5 or GPT-5.5.
Tag: MiniMax
MiniMax M3 Achieves 9x Faster AI Prefilling and 15x Faster Decoding for Long Contexts
MiniMax’s sparse attention architecture for its M3 model delivers massive speed improvements for processing large documents and code repositories.
MiniMax M3 Model Cuts AI Per-Token Compute to One-Twentieth of Previous Systems
MiniMax unveiled its M3 model with a sparse attention architecture that dramatically reduces compute costs while handling up to one million tokens.
MiniMax M3 Introduces Sparse Attention Architecture for Efficient AI Inference
AI company MiniMax released the M3 model featuring a new sparse attention design that significantly reduces computational requirements for inference tasks.