glm-5.3-flash API on AvalAI
glm-5.3-flashglm-5.3-flash is an AI model from Zai that is available through AvalAI's OpenAI-compatible API.
Use glm-5.3-flash with the AvalAI API
Keep the API key in an environment variable and send requests from server-side code.
curl https://api.avalai.ir/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AVALAI_API_KEY" \
-d '{"model": "glm-5.3-flash", "messages": [{"role": "user", "content": "Give a concise, practical solution."}]}'import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AVALAI_API_KEY"],
base_url="https://api.avalai.ir/v1",
)
response = client.chat.completions.create(
model="glm-5.3-flash",
messages=[{"role": "user", "content": "Give a concise, practical solution."}],
)
print(response.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.AVALAI_API_KEY,
baseURL: "https://api.avalai.ir/v1",
});
const response = await client.chat.completions.create({
model: "glm-5.3-flash",
messages: [{ role: "user", content: "Give a concise, practical solution." }],
});
console.log(response.choices[0].message.content);Create an API keyRead the quickstart
Capabilities and endpoints
- Function calling
- Prompt caching
- Reasoning
- Tool choice
glm-5.3-flash rate limits
| Tier | RPM | TPM |
|---|---|---|
| Basic | 3 | 40,000 |
| Tier 1 | 25 | 10,000,000 |
| Tier 2 | 250 | 4,000,000 |
| Tier 3 | 500 | 8,000,000 |
| Tier 4 | 750 | 10,000,000 |
| Tier 5 | 1,500 | 30,000,000 |
Related models
Frequently asked questions
What is the glm-5.3-flash API on AvalAI?
Create an AvalAI API key and send the model ID glm-5.3-flash to one of the supported endpoints shown on this page.
How much does the glm-5.3-flash API cost on AvalAI?
The glm-5.3-flash API on AvalAI costs $0.15 per 1M input tokens and $0.50 per 1M output tokens.
What is the token cost for glm-5.3-flash?
Current AvalAI pricing for glm-5.3-flash is $0.15 / 1M tokens for input and $0.50 / 1M tokens for output.
What is the glm-5.3-flash context window?
The recorded maximum input context for glm-5.3-flash is 991,000 tokens.
What are the rate limits for glm-5.3-flash on AvalAI?
On tier 5, glm-5.3-flash on AvalAI supports up to 1,500 requests per minute and 30,000,000 tokens per minute.
What features does glm-5.3-flash support on AvalAI?
glm-5.3-flash on AvalAI supports Function calling, Prompt caching, Reasoning, Tool choice.
What account tier is required to use glm-5.3-flash on AvalAI?
Calling glm-5.3-flash on AvalAI requires the Basic tier (tier 0) or higher.
Third-party data and freshness
- Pricing, availability, endpoints, and rate limits: AvalAI structured data (authoritative)