llama-4-maverick-17b-128e-instruct-fp8 API on AvalAI
llama-4-maverick-17b-128e-instruct-fp8llama-4-maverick-17b-128e-instruct-fp8 is an AI model from Meta that is available through AvalAI's OpenAI-compatible API.
Use llama-4-maverick-17b-128e-instruct-fp8 with the AvalAI API
Keep the API key in an environment variable and send requests from server-side code.
curl https://api.avalai.ir/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AVALAI_API_KEY" \
-d '{"model": "llama-4-maverick-17b-128e-instruct-fp8", "messages": [{"role": "user", "content": "Give a concise, practical solution."}]}'import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AVALAI_API_KEY"],
base_url="https://api.avalai.ir/v1",
)
response = client.chat.completions.create(
model="llama-4-maverick-17b-128e-instruct-fp8",
messages=[{"role": "user", "content": "Give a concise, practical solution."}],
)
print(response.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.AVALAI_API_KEY,
baseURL: "https://api.avalai.ir/v1",
});
const response = await client.chat.completions.create({
model: "llama-4-maverick-17b-128e-instruct-fp8",
messages: [{ role: "user", content: "Give a concise, practical solution." }],
});
console.log(response.choices[0].message.content);Create an API keyRead the quickstart
Capabilities and endpoints
- Function calling
- Tool choice
- Vision
llama-4-maverick-17b-128e-instruct-fp8 rate limits
| Tier | RPM | TPM |
|---|---|---|
| Tier 1 | 50 | 400,000 |
| Tier 2 | 150 | 1,000,000 |
| Tier 3 | 250 | 2,000,000 |
| Tier 4 | 500 | 4,000,000 |
| Tier 5 | 1,500 | 10,000,000 |
Related models
- llama-4-scout-17b-16e-instruct
- groq.llama-prompt-guard-2-86m
- groq.llama-prompt-guard-2-22m
- Llama Guard 4 12B
Frequently asked questions
What is the llama-4-maverick-17b-128e-instruct-fp8 API on AvalAI?
Create an AvalAI API key and send the model ID llama-4-maverick-17b-128e-instruct-fp8 to one of the supported endpoints shown on this page.
How much does the llama-4-maverick-17b-128e-instruct-fp8 API cost on AvalAI?
The llama-4-maverick-17b-128e-instruct-fp8 API on AvalAI costs $1.41 per 1M input tokens and $0.35 per 1M output tokens.
What is the token cost for llama-4-maverick-17b-128e-instruct-fp8?
Current AvalAI pricing for llama-4-maverick-17b-128e-instruct-fp8 is $1.41 / 1M tokens for input and $0.35 / 1M tokens for output.
What is the llama-4-maverick-17b-128e-instruct-fp8 context window?
The recorded maximum input context for llama-4-maverick-17b-128e-instruct-fp8 is 1,000,000 tokens.
What are the rate limits for llama-4-maverick-17b-128e-instruct-fp8 on AvalAI?
On tier 5, llama-4-maverick-17b-128e-instruct-fp8 on AvalAI supports up to 1,500 requests per minute and 10,000,000 tokens per minute.
What features does llama-4-maverick-17b-128e-instruct-fp8 support on AvalAI?
llama-4-maverick-17b-128e-instruct-fp8 on AvalAI supports Function calling, Tool choice, Vision.
What account tier is required to use llama-4-maverick-17b-128e-instruct-fp8 on AvalAI?
Calling llama-4-maverick-17b-128e-instruct-fp8 on AvalAI requires tier 1 or higher.
Third-party data and freshness
- Pricing, availability, endpoints, and rate limits: AvalAI structured data (authoritative)