Flash is not just a smaller GLM-5.3 endpoint
GLM-5.3-Flash uses a different 320B/18B-active hybrid architecture and adds native multimodality. Z.ai says sparse plus linear attention cuts attention compute and KV-cache requirements substantially versus the flagship. That design choice shows up in the live provider data: Flash is dramatically faster and cheaper while remaining in the same broad frontier band.
This is a separate deployment decision, not an effort setting. Use the flagship for the hardest text-only coding; use Flash when the system needs to look, act, and repeat economically.