MiniMax unveiled its M3 model with a sparse attention architecture that dramatically reduces compute costs while handling up to one million tokens.