- Good as a general-purpose model for text generation and everyday reasoning.
- Good when you want package quota first, with pay-as-you-go enabled by key configuration.
Zhipu active
GLM-5.3-flash API Pricing and Package Multipliers
GLM-5.3-flash
Model
glm-5.3-flash GLM-5.3-flash API Pricing
Official base: Input 0.15 / Output 0.5 / Cache read 0.03 / Cache write 0
| Channel | SU8 price | Billing unit | Quota multiplier | Effective multiplier |
|---|---|---|---|---|
| Zhipu | Input $0.12 Output $0.4 Cache read $0.024 Cache write $0 | USD / 1M | 5.376x | 0.8x |
Cache hits are billed at the cache read price. Cache writes use the cache write price, and the Console usage detail shows the final cost breakdown.
GLM-5.3-flash available packages and multipliers
| Package | Channel | SU8 price | Billing unit | Quota multiplier | Effective multiplier |
|---|---|---|---|---|---|
| Plus | Zhipu | Input $0.1375 Output $0.4585 Cache read $0.0275 Cache write $0 | USD / 1M | 42.3191x | 0.9169x |
| Max | Zhipu | Input $0.1101 Output $0.3668 Cache read $0.022 Cache write $0 | USD / 1M | 42.3191x | 0.7337x |
| Super Ultra | Zhipu | Input $0.0897 Output $0.2991 Cache read $0.0179 Cache write $0 | USD / 1M | 42.3191x | 0.5982x |
| Lite | Zhipu | Input $0.1529 Output $0.5096 Cache read $0.0306 Cache write $0 | USD / 1M | 42.3191x | 1.0192x |
| Pro | Zhipu | Input $0.1284 Output $0.4279 Cache read $0.0257 Cache write $0 | USD / 1M | 42.3191x | 0.8559x |
| Ultra | Zhipu | Input $0.0917 Output $0.3057 Cache read $0.0183 Cache write $0 | USD / 1M | 42.3191x | 0.6114x |
| SU500 | Zhipu | Input $0.0846 Output $0.2821 Cache read $0.0169 Cache write $0 | USD / 1M | 42.3191x | 0.5643x |
| SU750 | Zhipu | Input $0.0835 Output $0.2784 Cache read $0.0167 Cache write $0 | USD / 1M | 42.3191x | 0.5568x |
| SU1000 | Zhipu | Input $0.0846 Output $0.2821 Cache read $0.0169 Cache write $0 | USD / 1M | 42.3191x | 0.5643x |
GLM-5.3-flash OpenAI-compatible API access
After creating a key in the console, call this model with the unified model name glm-5.3-flash. When package and pay-as-you-go balance are both available, package quota is used first.
model: "glm-5.3-flash"
- Not ideal for legacy integrations that only accept OpenAI-compatible APIs.
- Not ideal when you need dedicated SLA terms, private capacity, or a fixed-cost enterprise commitment.
This model is available through the public catalog and unified gateway pricing configuration.
Open this model in the console
The public model page is for discovery and comparison. Actual API access, key creation, and model scopes are managed after sign-in.
