Reasoning Models
Explore advanced reasoning and problem-solving models available through AvalAI.
Introduction
Reasoning models are large language models trained to spend more compute on complex reasoning before they answer. Treat their reasoning as an internal process: ask for final answers, concise rationales, citations, or checklists, but do not ask the model to reveal hidden chain-of-thought. They excel in complex problem solving, coding, scientific reasoning, and multi-step planning for agentic workflows.
AvalAI offers access to several reasoning-capable models from different providers:
OpenAI Models:
gpt-5.6-sol: OpenAI's newest highest-capability GPT-5.6 flagship for hard agentic coding, knowledge work, scientific reasoning, and tool orchestration, 1M context windowgpt-5.6-terra: Balanced GPT-5.6 model for everyday production reasoning and agentic workflows, 1M context windowgpt-5.6-luna: Cost-efficient GPT-5.6 model for high-volume reasoning, support, and document workflows, 1M context windowgpt-5.5: Previous OpenAI flagship model with state-of-the-art reasoning across agentic coding, knowledge work, computer use, and scientific research (configurable effort: none/low/medium/high/xhigh), 1M context windowgpt-5.4-pro: OpenAI's highest-reasoning GPT-5.4 model for complex professional work, 1.05M context window, effort levels medium/high/xhighgpt-5.4: Frontier reasoning and agentic workflow model with configurable effort (none/low/medium/high/xhigh), 1.05M context windowgpt-5.4-mini: Fast, cost-efficient model with reasoning support (none, low, medium effort levels), 400K contextgpt-5.4-nano: Fastest and most affordable model with basic reasoning (none, low effort levels), 400K contextgpt-5-pro: OpenAI's advanced reasoning model with extended thinking capabilities for expert-level problem solving (Tier 2+ users, Responses API only)gpt-5.3-codex: OpenAI's most capable agentic coding model with reasoning tokens support (Responses API only)o1-pro: The most capable traditional reasoning model, but also the most expensiveo4-mini: Enhanced reasoning with improved efficiencyo3: Balanced reasoning capabilities with good performanceo3-mini: A smaller, faster model, generally less expensive per token
Google's Gemini Models:
gemini-3.6-flash: Google's July 2026 workhorse Flash reasoning model for agentic coding, knowledge work, multimodal analysis, efficient tool use, and 1M contextgemini-3.5-flash-lite: Google's fastest and most cost-effective Gemini 3.5-class reasoning model for high-throughput subagents, document parsing, extraction, and low-latency workloadsgemini-3.5-flash: Google's May 2026 Flash reasoning model with configurable thinking levels, strong coding and agentic tool-use performance, multimodal input, and 1M contextgemini-3.1-pro-preview: Google's advanced Pro-class model (Feb 2026) with natively multimodal reasoning, strong agentic performance, advanced coding, and long-context understandinggemini-3.1-flash-lite: Stable cost-efficient reasoning model optimized for high-frequency, lightweight agentic tasks with extremely low latencygemini-3.1-flash-lite-preview: Preview alias for Gemini 3.1 Flash-Lite with the same pricing and capabilitiesgemini-2.5-pro: Features enhanced thinking and reasoning capabilitiesgemini-2.5-flash: Google's first hybrid reasoning model with configurable thinking budgets
Anthropic Models:
claude-opus-5: Anthropic's new flagship Opus model with a 1M-token input window, 128K output capacity, adaptive thinking, configurable effort, stronger verification, and improved long-horizon coding, knowledge work, computer use, and scientific analysisclaude-opus-4-8: Previous Opus flagship with adaptive thinking, 5 effort levels (low/medium/high/xhigh/max), defaulthigheffort, mid-conversation system messages, and strong long-horizon agentic codingclaude-opus-4-7: Previous flagship with adaptive thinking, 5 effort levels (low/medium/high/xhigh/max), and task budgetsclaude-sonnet-5: Anthropic's most agentic Sonnet model with adaptive thinking and configurable effort levels; higher effort can match Opus 4.8 on some tasks at lower Sonnet-tier pricesclaude-sonnet-4-6: Strong reasoning capabilities with better efficiencyclaude-haiku-4-5: Fast, efficient reasoning for cost-sensitive deployments
Moonshot AI Models:
kimi-k3: Moonshot AI's new 2.8T-parameter flagship with a 1M-token context window, native vision, always-on reasoning, long-horizon coding, structured output, and tool callingkimi-latest: Stable alias that now resolves tokimi-k3with the same capabilities and pricing
For Kimi K3, use the top-level reasoning_effort: "max" field. K3 currently supports only the max effort level and keeps thinking enabled. Do not reuse the older K2.x thinking parameter or send fixed sampling fields such as temperature and top_p. In multi-turn and tool workflows, append the complete assistant message so the reasoning and tool-call context remains intact.
DeepSeek Models:
deepseek-v4-pro: DeepSeek's flagship reasoning model (1.6T total / 49B active params) with open-source SOTA Agentic Coding, 1M context window, provider-exposedreasoning_content,reasoning_effort: "high"/"max"deepseek-v4-flash: Stable ID now automatically backed by the official DeepSeek-V4-Flash-0731 release (284B total / 13B active params), with a 1M context window, stronger long-horizon coding and tool use, andreasoning_effort: "low"/"high"/"max"; no code or pricing change is requireddeepseek-reasoner: Legacy alias, now routes todeepseek-v4-pro(retiring July 24, 2026)deepseek-chat: Legacy alias, now routes todeepseek-v4-flash(retiring July 24, 2026)
XAI Models:
grok-4.5: XAI's new flagship model for coding, agentic tasks, and knowledge work with 1M context, fast serving, Chat Completions support, and partial Responses supportgrok-4.3: XAI's flagship reasoning model with 1M context window, function calling, structured outputs, and context-aware pricing above 200K tokensgrok-4.20-reasoning: Stable release with industry-leading speed and built-in reasoning, 2M context window, lowest hallucination rategrok-4.20-non-reasoning: Stable non-reasoning variant for tasks not requiring extended internal reasoning, 2M context window
MiniMax Models:
minimax-m3: New flagship model with frontier coding and agentic capabilities, 1M context window (MSA architecture), native multimodal input, and toggleable thinkingminimax-m2.7: Revolutionary self-evolution reasoning model (first model to deeply participate in its own evolution), 56.22% SWE-Pro, Agent Teams supportminimax-m2.7-highspeed: Ultra-fast self-evolution variant (~100 tokens per second output speed)minimax-m2.5: Earlier flagship model with SOTA coding (80.2% SWE-Bench Verified) and real-world productivityminimax-m2.5-lightning: Ultra-fast reasoning variant (~100 tokens per second output speed)minimax-m2.1: Flagship reasoning model with o3-level performance and 20x efficiencyminimax-m2.1-lightning: Fast reasoning variant (~100 tokens per second output speed)minimax-m2: Versatile reasoning model with strong general capabilities
Z.AI Models:
glm-5.2: Latest flagship reasoning model with 1M context, frontier coding, and long-horizon agentic engineeringglm-5.1: State-of-the-art reasoning with 58.4% SWE-Bench Pro, long-horizon optimization for agentic engineeringglm-5v-turbo: Multimodal reasoning and vision understanding with high-throughput visual processingglm-5-turbo: OpenClaw-optimized reasoning for tool invocation, persistent tasks, and long-chain executionglm-5: Flagship model with SOTA coding and agentic engineering capabilities
Alibaba Models:
qwen3.8-max: Newest 2.4T-parameter flagship for long-horizon coding, professional work, multimodal reasoning, and agentic planning, with hybrid thinking throughenable_thinking, a 1M context window, and up to 128K outputqwen3.7-max: Previous flagship agent-foundation model, hybrid thinking viaenable_thinking, 92.4 GPQA Diamond, 97.1 HMMT, 1M contextqwen3-max: Flagship Qwen3 Max model for complex reasoning and agentic workflows, hybrid thinking viaenable_thinking, 262K contextqwen3.6-plus: Agentic coding model with 78.8% SWE-bench Verified, 1M context default, frontier web developmentqwen3.6-flash: Fast hybrid-thinking model withenable_thinkingtoggle, 1M context windowqwen3.6-max-preview: Flagship preview with hybrid thinking, most capable Qwen3.6 modelqwen3.6-35b-a3b: Open-weight MoE (35B total/3B active) with thinking mode, 256K contextqwen3.6-27b: Dense vision-language reasoning model with 256K contextqwen3-coder-next: 80B MoE model (10B active) optimized for coding agents with 1M context
Fireworks.ai Models:
nemotron-3-ultra: NVIDIA's flagship large-scale Nemotron model for complex reasoning and agentic workflows, served via Fireworks.ai
Note: Some advanced models like o1-pro might have unique features and specific API endpoints (e.g., AvalAI's equivalent of the Responses API). Refer to the specific model documentation and the AvalAI API Reference for details.
When to use reasoning models
Reasoning models are best for tasks where correctness depends on planning, ambiguity resolution, or careful trade-offs. Use them when you need:
- Complex problem solving: math, science, finance, legal, policy, or engineering decisions with many constraints.
- Long-context synthesis: finding relationships across contracts, reports, transcripts, tickets, or retrieved documents.
- Agentic planning: deciding which tools to call, decomposing a workflow, or assigning steps to faster execution models.
- Code review and debugging: inspecting multi-file diffs, root-causing failures, or validating generated patches.
- Evaluation: judging model answers against nuanced rubrics or gold-standard requirements.
Think of reasoning models as planners and faster GPT-style models as workhorses. Use the planner for ambiguity, tool strategy, policy interpretation, long-context synthesis, or final validation. Use the workhorse for extraction, rewriting, classification, formatting, and other well-defined execution steps where speed and cost matter more than deep reasoning. Many production systems pair both: a reasoning model plans or validates, while a lower-latency model executes simple steps.
Prompting reasoning models effectively
Reasoning models need a different prompt style from classic GPT-style models:
- Keep the prompt simple and direct; start zero-shot before adding examples.
- Use a
developermessage for application rules and ausermessage for the task. - Use Markdown headings, XML tags, or other delimiters to separate rules, context, and examples.
- Be specific about constraints, success criteria, tools available, and what the final answer should contain.
- Avoid asking for hidden chain-of-thought or "think step by step"; ask for a short rationale, answer checklist, or cited evidence instead.
- If a reasoning-model snapshot suppresses Markdown but your UI needs Markdown, begin the developer message with
Formatting re-enabled.
Provider features vary. OpenAI models use reasoning.effort on /v1/responses; Kimi K3 uses top-level reasoning_effort: "max"; DeepSeek-V4-Flash-0731 accepts top-level reasoning_effort values low, high, and max; Alibaba Qwen models such as qwen3.8-max use enable_thinking with route-specific streaming requirements. Check the provider section for the selected model before copying parameters across families.
Get Started with Reasoning
For OpenAI reasoning models available through AvalAI, use /v1/responses when the model supports it. Responses preserves typed output items, supports reasoning.effort, and can keep relevant reasoning items across tool calls with previous_response_id or by replaying prior output items. Keep Chat Completions examples for provider-specific routes or legacy integrations that do not expose Responses yet.
Responses-first reasoning checklist
- Start with
gpt-5.6-terraorgpt-5.6-solandreasoning: {"effort": "medium"}for balanced quality, latency, and cost; trylowfor faster support or drafting flows andhigh/xhighonly when evals justify the added tokens. - Reserve enough
max_output_tokensfor both visible output and hidden reasoning tokens. If a response returnsstatus: "incomplete"withincomplete_details.reason: "max_output_tokens", the model may have spent the entire budget on reasoning and produced no visible text. Increase the budget, lowerreasoning.efforttolowornonewhen supported, or simplify the task. - When a reasoning model calls tools, set
store: truewhen your retention policy allows it, then pass back the reasoning items and function-call items from the prior response, or continue withprevious_response_id. - For stateless or zero-retention style flows, include
reasoning.encrypted_contentwhen supported so encrypted reasoning items can be round-tripped. - Request
reasoning.summaryonly when the selected model and account support summaries; summaries are for observability, not raw chain-of-thought. - Prompt reasoning models with goals, constraints, available tools, and success criteria. Avoid asking for hidden chain-of-thought; ask for a concise final explanation or checklist instead.
Latency and visible progress
For hard tasks, raising reasoning.effort can improve answer quality but also delays the first useful token. OpenAI's reasoning guidance recommends asking the model for a short preamble when you need faster visible progress. In AvalAI, use this pattern for streamed developer tools, review assistants, and long-running agentic workflows:
- Ask for one short progress sentence before deeper analysis, not hidden chain-of-thought.
- Render the preamble as status text; do not treat it as the final answer.
- Keep the final answer contract explicit: decision, evidence, checklist, patch plan, or next action.
- If the route supports
phase, preservephase: "commentary"preambles andphase: "final_answer"final responses when replaying state.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AVALAI_API_KEY"],
base_url="https://api.avalai.ir/v1",
)
response = client.responses.create(
model="gpt-5.6-sol",
reasoning={"effort": "high"},
instructions=(
"For difficult work, first give one short progress sentence. "
"Do not reveal hidden chain-of-thought. Finish with a concise "
"decision, evidence, and verification checklist."
),
input="Review this migration plan for production risks: ...",
)
print(response.output_text)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.AVALAI_API_KEY,
baseURL: "https://api.avalai.ir/v1",
});
const response = await client.responses.create({
model: "gpt-5.6-sol",
reasoning: { effort: "high" },
instructions:
"For difficult work, first give one short progress sentence. Do not reveal hidden chain-of-thought. Finish with a concise decision, evidence, and verification checklist.",
input: "Review this migration plan for production risks: ...",
});
console.log(response.output_text);Long-running tool flows and phase
OpenAI's current Responses guidance recommends preserving the assistant phase field for long-running or tool-heavy GPT-5.5/GPT-5.4 workflows. Treat this as model- and route-dependent in AvalAI: prefer previous_response_id when storage is allowed, and if you manually replay assistant history, keep any original phase values unchanged.
Use phase: "commentary" for intermediate assistant updates before tool calls, and phase: "final_answer" for the completed answer. Do not add phase to user messages. Dropping phase metadata can make an intermediate note look like the final answer in multi-step flows.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AVALAI_API_KEY"],
base_url="https://api.avalai.ir/v1",
)
response = client.responses.create(
model="gpt-5.6-sol",
input=[
{
"role": "assistant",
"phase": "commentary",
"content": "I will inspect the logs before recommending a fix.",
},
{
"role": "assistant",
"phase": "final_answer",
"content": "Root cause: cache invalidation race.",
},
{"role": "user", "content": "Now give me a rollout-safe fix plan."},
],
)
print(response.output_text)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.AVALAI_API_KEY,
baseURL: "https://api.avalai.ir/v1",
});
const response = await client.responses.create({
model: "gpt-5.6-sol",
input: [
{
role: "assistant",
phase: "commentary",
content: "I will inspect the logs before recommending a fix.",
},
{
role: "assistant",
phase: "final_answer",
content: "Root cause: cache invalidation race.",
},
{ role: "user", content: "Now give me a rollout-safe fix plan." },
],
});
console.log(response.output_text);Example: Using a reasoning model
PROMPT='Write a bash script that takes a matrix represented as a string with format '\''[1,2],[3,4],[5,6]'\'' and prints the transpose in the same format.'
curl https://api.avalai.ir/v1/responses \
-H "Authorization: Bearer $AVALAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-sol",
"reasoning": {"effort": "medium"},
"input": [
{
"role": "user",
"content": "'"$PROMPT"'"
}
]
}'import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AVALAI_API_KEY"],
base_url="https://api.avalai.ir/v1",
)
prompt = """
Write a bash script that takes a matrix represented as a string with
format '[1,2],[3,4],[5,6]' and prints the transpose in the same format.
"""
try:
response = client.responses.create(
model="gpt-5.6-sol",
reasoning={"effort": "medium"},
input=[{"role": "user", "content": prompt}],
)
print(response.output_text)
except Exception as e:
print(f"An API error occurred: {e}")import OpenAI from "openai"; // Use the standard OpenAI library configured for AvalAI
import * as dotenv from "dotenv";
dotenv.config();
const client = new OpenAI({
apiKey: process.env.AVALAI_API_KEY,
baseURL: "https://api.avalai.ir/v1", // AvalAI endpoint
});
const prompt = `
Write a bash script that takes a matrix represented as a string with
format '[1,2],[3,4],[5,6]' and prints the transpose in the same format.
`;
async function runReasoning() {
try {
const response = await client.responses.create({
model: "gpt-5.6-sol",
reasoning: { effort: "medium" },
input: [{ role: "user", content: prompt }],
});
console.log(response.output_text);
} catch (error) {
console.error("An API error occurred:", error);
}
}
runReasoning();package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
)
func main() {
prompt := `
Write a bash script that takes a matrix represented as a string with
format '[1,2],[3,4],[5,6]' and prints the transpose in the same format.
`
payload := map[string]any{
"model": "gpt-5.5",
"reasoning": map[string]string{"effort": "medium"},
"input": []map[string]string{
{
"role": "user",
"content": prompt,
},
},
}
body, err := json.Marshal(payload)
if err != nil {
panic(err)
}
req, err := http.NewRequest("POST", "https://api.avalai.ir/v1/responses", bytes.NewBuffer(body))
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+os.Getenv("AVALAI_API_KEY"))
req.Header.Set("Content-Type", "application/json")
resp, err := http.DefaultClient.Do(req)
if err != nil {
fmt.Printf("Responses API error: %v\n", err)
return
}
defer resp.Body.Close()
responseBody, err := io.ReadAll(resp.Body)
if err != nil {
panic(err)
}
fmt.Println(string(responseBody))
}<?php
$apiKey = getenv('AVALAI_API_KEY');
$prompt = <<<PROMPT
Write a bash script that takes a matrix represented as a string with
format '[1,2],[3,4],[5,6]' and prints the transpose in the same format.
PROMPT;
$payload = [
'model' => 'gpt-5.5',
'reasoning' => ['effort' => 'medium'],
'input' => [
['role' => 'user', 'content' => $prompt],
],
];
$ch = curl_init('https://api.avalai.ir/v1/responses');
curl_setopt_array($ch, [
CURLOPT_RETURNTRANSFER => true,
CURLOPT_POST => true,
CURLOPT_HTTPHEADER => [
'Authorization: Bearer ' . $apiKey,
'Content-Type: application/json',
],
CURLOPT_POSTFIELDS => json_encode($payload),
]);
$response = curl_exec($ch);
curl_close($ch);
echo $response;
?>Reasoning Effort
Some reasoning models accept parameters like reasoning.effort to guide the amount of internal reasoning performed before generating a response. Potential values include:
none: Disables explicit reasoning where supported, favoring speed.minimal: Uses the smallest supported reasoning budget when a model exposes this value.low: Favors speed and lower token usage.medium(often default): Balances speed and reasoning quality.high: Favors more thorough reasoning, potentially using more tokens and taking longer.xhighormax: Uses the most reasoning for the hardest coding, math, and agentic tasks where supported.
Consult the specific model documentation on AvalAI for supported parameters and their effects.
Choosing effort in production
Treat effort as a latency/cost/quality tuning knob, not as the first fix for a weak prompt. Start with the lowest setting that passes your evals, then raise it only for tasks where the extra reasoning tokens produce measurable quality gains.
| Workload | Suggested starting point | Why |
|---|---|---|
| Voice, classification, simple retrieval | none or low when supported | Prioritizes first-token latency and lower token use. |
| Customer support, drafting, tool planning | low or medium | Leaves room for tool strategy without over-spending. |
| Coding, research, spreadsheet/document analysis | medium | Matches the balanced default for current OpenAI reasoning guidance such as gpt-5.5. |
| Deep research, security review, hard debugging | high or xhigh after evals | Useful when accuracy matters more than latency and cost. |
Defaults are model-dependent. Do not assume that medium is universal across OpenAI, Anthropic, Gemini, DeepSeek, XAI, or other providers; record the chosen effort in your eval traces and compare quality, latency, output_tokens, and output_tokens_details.reasoning_tokens before changing production defaults.
How Reasoning Works
Reasoning models utilize reasoning tokens internally, in addition to the standard input and output tokens. These tokens represent the model's "thought process" – breaking down the prompt, exploring approaches, and planning the response. After this internal reasoning phase, the model generates the final visible output tokens and typically discards the reasoning tokens from its context memory for subsequent turns.
(Diagram Source: OpenAI)
Important: Reasoning tokens can be hidden from the final response text and still be billable and occupy the model's context window during generation. Across providers and models, they are billed at the selected model's output-token rate. When output_tokens already includes them, output_tokens_details.reasoning_tokens is a breakdown, so adding reasoning_tokens to output_tokens would double-count them. If a route reports visible output and reasoning separately, apply the same output-token rate to both counts.
Provider-Specific Reasoning Settings
Different model providers implement reasoning capabilities in their own ways. Here's how to use reasoning features with each provider through AvalAI:
Gemini Models Reasoning Settings
Google's current Gemini 3.6, 3.5, and 3.1 reasoning-capable models support configurable thinking through Gemini generationConfig.thinkingConfig. For gemini-3.6-flash, gemini-3.5-flash-lite, gemini-3.5-flash, gemini-3.1-pro-preview, gemini-3.1-flash-lite, and gemini-3.1-flash-lite-preview, use thinkingLevel to balance quality, latency, and cost.
When using OpenAI client libraries with AvalAI, pass Gemini-specific settings through the extra_body parameter:
# Python example - Gemini thinking levels
response = client.chat.completions.create(
model="gemini-3.6-flash",
messages=[
{
"role": "user",
"content": "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ...",
}
],
extra_body={"generationConfig": {"thinkingConfig": {"thinkingLevel": "high"}}},
)// JavaScript example - Gemini thinking levels
const response = await client.chat.completions.create({
model: "gemini-3.6-flash",
messages: [
{
role: "user",
content: "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ...",
},
],
// @ts-expect-error extra_body is supported by AvalAI for Gemini-specific options
extra_body: {
generationConfig: {
thinkingConfig: { thinkingLevel: "high" },
},
},
});Responses API version
Gemini 3.6 Flash and Gemini 3.5 Flash-Lite have partial /v1/responses support. Use this version when the parameters and tools required by your workflow are supported; messages moves to input, and the final text is read from response.output_text.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AVALAI_API_KEY"],
base_url="https://api.avalai.ir/v1",
)
response = client.responses.create(
model="gemini-3.6-flash",
instructions="You are a helpful assistant.",
input="Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ...",
)
print(response.output_text)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.AVALAI_API_KEY,
baseURL: "https://api.avalai.ir/v1",
});
const response = await client.responses.create({
model: "gemini-3.6-flash",
instructions: "You are a helpful assistant.",
input: "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ...",
});
console.log(response.output_text);curl https://api.avalai.ir/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AVALAI_API_KEY" \
-d '
{
"model": "gemini-3.6-flash",
"input": "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ...",
"instructions": "You are a helpful assistant."
}'messages→input- system message →
instructionsor adeveloperitem choices[0].message.content→response.output_text- for tools and multimodal output, inspect
response.outputby itemtype.
The Gemini 3.6/3.5/3.1 thinking settings include:
thinkingLevel: Controls reasoning depth. Supported levels depend on the selected model and current route.- Higher thinking levels generally improve multi-step reasoning and tool-use quality, but can increase latency and token usage.
- Use a mid-range supported level for balanced production defaults and a higher supported level for complex coding, agentic, or analytical tasks.
For the native Gemini v1beta endpoint, pass the same configuration directly in the request body:
curl https://api.avalai.ir/v1beta/models/gemini-3.6-flash:generateContent \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AVALAI_API_KEY" \
-d '{
"contents": [{"role": "user", "parts": [{"text": "Plan a resilient migration from a monolith to microservices."}]}],
"generationConfig": {
"thinkingConfig": {"thinkingLevel": "high"}
}
}'Legacy note
thinking_budget is specific to Gemini 2.5 Flash. For the most current information, reference the official Google AI documentation.
Gemini 2.5 Flash uses budget-based thinking controls:
# Python example - Gemini 2.5 Flash thinking budget
response = client.chat.completions.create(
model="gemini-2.5-flash",
messages=[
{
"role": "user",
"content": "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ...",
}
],
extra_body={"thinking": {"type": "enabled", "budget_tokens": 2000}},
)// JavaScript example - Gemini 2.5 Flash thinking budget
const responseAlt = await client.chat.completions.create({
model: "gemini-2.5-flash",
messages: [{ role: "user", content: "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ..." }],
// @ts-expect-error thinking is an undocumented provider-specific parameter
thinking: { type: "enabled", budget_tokens: 2000 },
});Responses API version This version uses `gpt-5.5` because `gemini-2.5-flash` may not be enabled for `/v1/responses` in the current AvalAI model data.
Use this version when the selected model supports /v1/responses. messages moves to input, and the final text is read from response.output_text.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AVALAI_API_KEY"],
base_url="https://api.avalai.ir/v1",
)
response = client.responses.create(
model="gpt-5.6-sol",
input=[
{
"role": "user",
"content": [
{
"type": "input_text",
"text": "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ...",
},
{"type": "input_file", "file_id": "file_abc123"},
],
}
],
)
print(response.output_text)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.AVALAI_API_KEY,
baseURL: "https://api.avalai.ir/v1",
});
const response = await client.responses.create({
model: "gpt-5.6-sol",
input: [
{
role: "user",
content: [
{ type: "input_text", text: "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ..." },
{ type: "input_file", file_id: "file_abc123" },
],
},
],
});
console.log(response.output_text);curl https://api.avalai.ir/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AVALAI_API_KEY" \
-d '
{
"model": "gpt-5.6-sol",
"input": [
{
"role": "user",
"content": [
{
"type": "input_text",
"text": "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ..."
},
{
"type": "input_file",
"file_id": "file_abc123"
}
]
}
]
}'messages→input- system message →
instructionsor adeveloperitem choices[0].message.content→response.output_text- for tools and multimodal output, inspect
response.outputby itemtype.
By setting appropriate thinking levels or budgets, you can control both the depth of reasoning and the cost of your API calls.
Anthropic Models Reasoning Settings
Anthropic's Claude models (specifically claude-opus-5, claude-sonnet-5, claude-opus-4-8, claude-opus-4-7, claude-sonnet-4-6, claude-haiku-4-5) support configurable reasoning through provider-specific thinking settings. For Claude Opus 5, use thinking: {"type": "adaptive"} together with output_config.effort; do not send a fixed extended-thinking budget. Start at medium or high, then increase to xhigh or max only when evaluations show a material improvement in task success.
response = client.chat.completions.create(
model="claude-opus-5",
messages=[
{
"role": "user",
"content": "Analyze this migration design, identify failure modes, and verify the rollback plan.",
}
],
extra_body={
"thinking": {"type": "adaptive"},
"output_config": {"effort": "high"},
},
)Responses API version Claude Opus 5 has partial `/v1/responses` support; verify required fields and tools before production use.
Use this version when the selected model supports /v1/responses. messages moves to input, and the final text is read from response.output_text.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AVALAI_API_KEY"],
base_url="https://api.avalai.ir/v1",
)
response = client.responses.create(
model="gpt-5.6-sol",
instructions="You are a helpful assistant.",
input="Analyze this migration design, identify failure modes, and verify the rollback plan.",
)
print(response.output_text)messages→input- system message →
instructionsor adeveloperitem choices[0].message.content→response.output_text- for tools and multimodal output, inspect
response.outputby itemtype.
Anthropic effort controls are provider-specific; do not copy Gemini thinking parameters or OpenAI reasoning.effort fields without checking the selected route. Request concise rationales or verification evidence rather than hidden chain-of-thought.
OpenAI Models Reasoning Settings
OpenAI's latest reasoning-capable models (gpt-5.5, gpt-5.4-pro, gpt-5.4, gpt-5.4-mini, gpt-5.4-nano, gpt-5.3-codex, gpt-5-pro, o4-mini, o3, and o3-mini) have built-in reasoning capabilities that are automatically engaged when needed. For these models, you may see reasoning tokens included in your usage statistics, but the reasoning process is more integrated into the model's operation.
For more advanced control, some endpoints may support parameters like reasoning.effort to guide the amount of internal reasoning:
response = client.chat.completions.create(
model="gpt-5.6-sol",
messages=[
{
"role": "user",
"content": "Design an algorithm to solve this optimization problem: ...",
}
],
extra_body={"reasoning": {"effort": "high"}}, # Request more thorough reasoning
)Responses API version
Use this version when the selected model supports /v1/responses. messages moves to input, and the final text is read from response.output_text.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AVALAI_API_KEY"],
base_url="https://api.avalai.ir/v1",
)
response = client.responses.create(
model="gpt-5.6-sol",
instructions="You are a helpful assistant.",
input="Design an algorithm to solve this optimization problem: ...",
)
print(response.output_text)messages→input- system message →
instructionsor adeveloperitem choices[0].message.content→response.output_text- for tools and multimodal output, inspect
response.outputby itemtype.
DeepSeek Models Reasoning Settings
DeepSeek offers reasoning capabilities through the V4 flagship family (deepseek-v4-pro and deepseek-v4-flash). The stable deepseek-v4-flash ID now uses the official DeepSeek-V4-Flash-0731 release automatically at the same price, with stronger agentic capabilities and low, high, and max reasoning-effort levels. In thinking mode, supported routes can return provider-specific reasoning_content for continuity/observability alongside the final answer. The legacy aliases deepseek-reasoner (→ deepseek-v4-pro) and deepseek-chat (→ deepseek-v4-flash) continue to work but will be retired on July 24, 2026.
Key Features:
reasoning_content: Contains the provider-exposed reasoning trace used for tool-call continuity and observabilitycontent: Contains the final answer- Reasoning trace: Use the provider-exposed trace for debugging, tool-call continuity, and observability
- Tool Call Integration: Thinking mode works with function calling
- Thinking Mode Toggle: Use
extra_body={"thinking": {"type": "enabled"}}or{"type": "disabled"}where the selected V4 route exposes the toggle - Reasoning Effort: DeepSeek-V4-Flash-0731 supports
reasoning_effort: "low","high", or"max"; verify V4-Pro values separately before sharing configuration across the family - Context Window: V4 flagships default to a 1M-token context window
Basic Usage (thinking mode, V4-Pro):
response = client.chat.completions.create(
model="deepseek-v4-pro",
messages=[
{
"role": "user",
"content": "9.11 and 9.8, which is greater?",
}
],
reasoning_effort="high",
extra_body={"thinking": {"type": "enabled"}},
)
# Access the reasoning process
reasoning = response.choices[0].message.reasoning_content
# Access the final answer
answer = response.choices[0].message.contentResponses API version This version uses `gpt-5.5` because `deepseek-v4-pro` may not be enabled for `/v1/responses` in the current AvalAI model data.
Use this version when the selected model supports /v1/responses. messages moves to input, and the final text is read from response.output_text.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AVALAI_API_KEY"],
base_url="https://api.avalai.ir/v1",
)
response = client.responses.create(
model="gpt-5.6-sol",
instructions="You are a helpful assistant.",
input="9.11 and 9.8, which is greater?",
)
print(response.output_text)messages→input- system message →
instructionsor adeveloperitem choices[0].message.content→response.output_text- for tools and multimodal output, inspect
response.outputby itemtype.
Fast, Economical Reasoning (V4-Flash):
response = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[
{
"role": "user",
"content": "Summarize the trade-offs between thinking and non-thinking modes.",
}
],
extra_body={"thinking": {"type": "enabled"}},
)Responses API version This version uses `gpt-5.5` because `deepseek-v4-flash` may not be enabled for `/v1/responses` in the current AvalAI model data.
Use this version when the selected model supports /v1/responses. messages moves to input, and the final text is read from response.output_text.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AVALAI_API_KEY"],
base_url="https://api.avalai.ir/v1",
)
response = client.responses.create(
model="gpt-5.6-sol",
instructions="You are a helpful assistant.",
input="Summarize the trade-offs between thinking and non-thinking modes.",
)
print(response.output_text)messages→input- system message →
instructionsor adeveloperitem choices[0].message.content→response.output_text- for tools and multimodal output, inspect
response.outputby itemtype.
⚠️ Critical: Tool Calls with Thinking Mode
When using tool calls with DeepSeek V4 models in thinking mode (including deepseek-v4-pro, deepseek-v4-flash with thinking enabled, or the legacy deepseek-reasoner alias), you must pass the reasoning_content back to the API in subsequent requests within the same turn. Failure to do so will result in an error:
Missing reasoning_content field in the assistant messageCorrect Tool Call Implementation:
# When the model returns tool_calls, include reasoning_content in the assistant message
assistant_message = {
"role": "assistant",
"content": message.content or "",
"tool_calls": [...],
"reasoning_content": message.reasoning_content, # CRITICAL: Must include this
}
messages.append(assistant_message)Multi-turn Conversation Rules:
- Within a single turn (while processing tool calls): Always include
reasoning_content - Between turns (new user message): Only pass
content, notreasoning_content
For complete documentation with examples in multiple languages, see the DeepSeek Models Documentation.
Official Reference: DeepSeek Thinking Mode - Tool Calls
Direct HTTP Requests
When making direct HTTP requests or using curl, you can include these parameters directly in the request body:
curl https://api.avalai.ir/v1/chat/completions \
-H "Authorization: Bearer $AVALAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.1-flash-lite-preview",
"messages": [{"role": "user", "content": "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ..."}],
"thinking": {"type": "enabled", "budget_tokens": 2000}
}'Responses API version This version uses `gpt-5.5` because `gemini-3.1-flash-lite-preview` may not be enabled for `/v1/responses` in the current AvalAI model data.
Use this version when the selected model supports /v1/responses. messages moves to input, and the final text is read from response.output_text.
curl https://api.avalai.ir/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AVALAI_API_KEY" \
-d '
{
"model": "gpt-5.6-sol",
"input": "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ...",
"instructions": "You are a helpful assistant."
}'messages→input- system message →
instructionsor adeveloperitem choices[0].message.content→response.output_text- for tools and multimodal output, inspect
response.outputby itemtype.
By adjusting these parameters, you can balance reasoning depth against cost and speed for your specific use case across different model providers.
Managing the Context Window
Ensure sufficient space remains in the context window for both the expected output and the internal reasoning tokens. Complex problems might require thousands or even tens of thousands of reasoning tokens.
You can often find the breakdown of token usage (including reasoning tokens, if exposed by the API) in the usage object of the API response. For OpenAI Responses, output_tokens is the generated-output total including reasoning tokens; the nested reasoning_tokens value is not an additional amount to add.
Example OpenAI Responses usage object:
{
"usage": {
"input_tokens": 75,
"input_tokens_details": {
"cached_tokens": 0
},
"output_tokens": 1186,
"output_tokens_details": {
"reasoning_tokens": 1024
},
"total_tokens": 1261
}
}Check the AvalAI Models documentation for context window lengths for specific models.
Use usage telemetry as a tuning loop:
- Log
model,reasoning.effort,max_output_tokens, latency, status, and token usage for each eval case. - Watch for high
reasoning_tokenswith low answer-quality gains; lower effort or simplify the prompt when this happens. - Watch for
status: "incomplete"or missing visible output; increasemax_output_tokens, split the task, or reduce retrieved context. - Keep separate baselines for Chat Completions and Responses, because Responses can preserve reasoning items across tool calls while Chat Completions remains stateless.
Controlling Costs
To manage costs:
- Be mindful of the
reasoning.effort(or equivalent) parameter if available. Higher effort usually means more tokens. - Use the
max_tokens(or equivalent parameter likemax_output_tokens) in your API request to limit the total number of tokens generated (reasoning + output).
Allocating Space for Reasoning
max_output_tokens (Responses), max_completion_tokens (Chat Completions), and legacy max_tokens are shared generation budgets on reasoning-capable models: hidden reasoning tokens and visible answer tokens can both consume them. They do not reserve a separate allowance for the final answer.
If reasoning consumes the entire limit, the response can contain a reasoning item but no text. This is a token-budget exhaustion case, not necessarily a content-filter or model-availability problem. Typical evidence includes:
- Responses:
status: "incomplete"andincomplete_details.reason: "max_output_tokens". - Usage:
output_tokens_details.reasoning_tokensis close to the totaloutput_tokens, while visible text is empty or missing. - Chat Completions:
finish_reason: "length", possibly before useful visible content is produced.
To recover, increase the applicable output limit within the model's supported maximum, lower reasoning.effort to low or none when the model supports it, simplify or split the task, and leave measured headroom for the final answer. Do not size the limit only from the desired visible answer length. Record usage by model and prompt class because reasoning demand varies between requests, even for the same model.
Handling Incomplete Responses (Example)
Your code should check for indicators that generation was stopped due to token limits.
import json
# ... (client setup and initial request as before) ...
try:
response = client.chat.completions.create(
model="gpt-5.6-sol",
messages=[{"role": "user", "content": prompt}],
max_tokens=300, # Limit total generated tokens (reasoning + output)
# Add reasoning parameters if applicable
)
finish_reason = response.choices[0].finish_reason
output_text = response.choices[0].message.content
if finish_reason == "length": # Standard finish reason for max_tokens
print("Ran out of tokens (max_tokens reached).")
if output_text:
print("Partial output:", output_text)
else:
# This implies the limit was hit during internal reasoning
print("Ran out of tokens during reasoning phase.")
elif finish_reason == "stop":
print("Completed successfully:")
print(output_text)
else:
print(f"Finished with reason: {finish_reason}")
if output_text:
print("Output:", output_text)
except Exception as e:
print(f"An API error occurred: {e}")// ... (client setup and initial request as before) ...
async function runReasoningWithLimit() {
try {
const response = await client.chat.completions.create({
model: "gpt-5.6-sol",
messages: [{ role: "user", content: prompt }],
max_tokens: 300, // Limit total generated tokens
// Add reasoning parameters if applicable
});
const finish_reason = response.choices[0].finish_reason;
const output_text = response.choices[0].message.content;
if (finish_reason === "length") {
// Standard finish reason for max_tokens
console.log("Ran out of tokens (max_tokens reached).");
if (output_text) {
console.log("Partial output:", output_text);
} else {
console.log("Ran out of tokens during reasoning phase.");
}
} else if (finish_reason === "stop") {
console.log("Completed successfully:");
console.log(output_text);
} else {
console.log(`Finished with reason: ${finish_reason}`);
if (output_text) {
console.log("Output:", output_text);
}
}
} catch (error) {
console.error("An API error occurred:", error);
}
}
runReasoningWithLimit();package main
import (
"context"
"fmt"
"os"
openai "github.com/openai/openai-go"
)
func main() {
// ... (client setup as before) ...
prompt := "..." // Your prompt here
maxTokens := 300 // Define max_tokens
resp, err := client.CreateChatCompletion(
context.Background(),
openai.ChatCompletionRequest{
Model: "gpt-5.5",
Messages: []openai.ChatCompletionMessage{
{Role: openai.ChatMessageRoleUser, Content: prompt},
},
MaxTokens: maxTokens, // Limit total generated tokens
// Add reasoning parameters if applicable
},
)
if err != nil {
fmt.Printf("ChatCompletion error: %v\n", err)
return
}
finishReason := resp.Choices[0].FinishReason
outputText := resp.Choices[0].Message.Content
if finishReason == openai.FinishReasonLength { // Check for length finish reason
fmt.Println("Ran out of tokens (max_tokens reached).")
if outputText != "" {
fmt.Println("Partial output:", outputText)
} else {
fmt.Println("Ran out of tokens during reasoning phase.")
}
} else if finishReason == openai.FinishReasonStop {
fmt.Println("Completed successfully:")
fmt.Println(outputText)
} else {
fmt.Printf("Finished with reason: %s\n", finishReason)
if outputText != "" {
fmt.Println("Output:", outputText)
}
}
}<?php
require 'vendor/autoload.php';
// ... (client setup as before) ...
$prompt = "..."; // Your prompt here
$maxTokens = 300; // Define max_tokens
try {
$response = $client->chat()->create([
'model' => 'gpt-5.5',
'messages' => [
['role' => 'user', 'content' => $prompt],
],
'max_tokens' => $maxTokens, // Limit total generated tokens
// Add reasoning parameters if applicable
]);
$finishReason = $response->choices[0]->finishReason;
// Ensure content exists before accessing
$outputText = $response->choices[0]->message->content ?? null;
if ($finishReason === 'length') { // Check for length finish reason
echo "Ran out of tokens (max_tokens reached).\n";
if ($outputText) {
echo "Partial output: " . $outputText . "\n";
} else {
echo "Ran out of tokens during reasoning phase.\n";
}
} elseif ($finishReason === 'stop') {
echo "Completed successfully:\n";
echo $outputText . "\n";
} else {
echo "Finished with reason: " . $finishReason . "\n";
if ($outputText) {
echo "Output: " . $outputText . "\n";
}
}
} catch (Exception $e) {
echo "An API error occurred: " . $e->getMessage() . "\n";
}
?>Responses API version
Use this version when the selected model supports /v1/responses. Chat Completions reports token exhaustion with finish_reason: "length"; Responses reports it with status: "incomplete" and incomplete_details.reason: "max_output_tokens".
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AVALAI_API_KEY"],
base_url="https://api.avalai.ir/v1",
)
prompt = """
Write a bash script that takes a matrix represented as a string with
format '[1,2],[3,4],[5,6]' and prints the transpose in the same format.
"""
response = client.responses.create(
model="gpt-5.6-sol",
reasoning={"effort": "medium"},
input=[{"role": "user", "content": prompt}],
max_output_tokens=300,
)
if (
response.status == "incomplete"
and response.incomplete_details.reason == "max_output_tokens"
):
print("Ran out of tokens.")
if response.output_text:
print("Partial output:", response.output_text)
else:
print("Ran out of tokens during reasoning.")
else:
print(response.output_text)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.AVALAI_API_KEY,
baseURL: "https://api.avalai.ir/v1",
});
const prompt = `
Write a bash script that takes a matrix represented as a string with
format '[1,2],[3,4],[5,6]' and prints the transpose in the same format.
`;
const response = await client.responses.create({
model: "gpt-5.6-sol",
reasoning: { effort: "medium" },
input: [{ role: "user", content: prompt }],
max_output_tokens: 300,
});
if (
response.status === "incomplete" &&
response.incomplete_details?.reason === "max_output_tokens"
) {
console.log("Ran out of tokens.");
if (response.output_text) {
console.log("Partial output:", response.output_text);
} else {
console.log("Ran out of tokens during reasoning.");
}
} else {
console.log(response.output_text);
}PROMPT='Write a bash script that takes a matrix represented as a string with format "[1,2],[3,4],[5,6]" and prints the transpose in the same format.'
jq -n --arg prompt "$PROMPT" '{
model: "gpt-5.6-sol",
reasoning: {effort: "medium"},
input: [{role: "user", content: $prompt}],
max_output_tokens: 300
}' | curl https://api.avalai.ir/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AVALAI_API_KEY" \
-d @- | jq '{
status,
incomplete_reason: .incomplete_details.reason,
output_text
}'messages→input- system message →
instructionsor adeveloperitem max_tokens→max_output_tokenschoices[0].finish_reason == "length"→status == "incomplete"withincomplete_details.reasonchoices[0].message.content→response.output_text- for tools and multimodal output, inspect
response.outputby itemtype.
Advice on Prompting
Prompting reasoning models can differ slightly from prompting standard GPT models.
- Reasoning Models: Often perform well with higher-level goals and fewer prescribed intermediate steps. Think of them as senior collaborators you can trust to figure out the details.
- GPT Models: Often benefit from very precise instructions and clear definitions of the desired output format. Think of them as junior collaborators needing explicit guidance.
Experiment with providing the overall objective and letting the reasoning model determine the best path, versus providing detailed steps.
Prompt Examples
(Note: The following examples use current reasoning-capable models such as gpt-5.5, deepseek-v4-pro, qwen3.7-max, and glm-5.2. Adjust parameters/endpoints as needed based on AvalAI's specific implementation.)
1. Coding (Refactoring)
Task: Refactor a React component to change text color based on data.
// --- Calling Code (Node.js) ---
import OpenAI from "openai";
import * as dotenv from "dotenv";
dotenv.config();
const client = new OpenAI({
apiKey: process.env.AVALAI_API_KEY,
baseURL: "https://api.avalai.ir/v1",
});
// Note the change here: Use indentation for the inner code example
// instead of triple backticks within the prompt string.
const prompt = `
Instructions:- Given the React component below, change it so that nonfiction books have red text.
- Return only the refactored React component code in your reply.
- Do not include explanations or markdown code blocks.
- Use four spaces for indentation.
- Keep lines under 80 columns.
Original Code:
const books = [
{ title: 'Dune', category: 'fiction', id: 1 },
{ title: 'Frankenstein', category: 'fiction', id: 2 },
{ title: 'Moneyball', category: 'nonfiction', id: 3 },
];
export default function BookList() {
const listItems = books.map(book =>
<li>
{book.title}
</li>
);
return (
<ul>{listItems}</ul>
);
}
`.trim();
async function refactorCode() {
try {
const response = await client.chat.completions.create({
model: "gpt-5.6-sol", // Use a suitable reasoning model from AvalAI
messages: [{ role: "user", content: prompt }],
temperature: 0.1, // Lower temperature for more predictable code output
});
console.log(response.choices[0].message.content);
} catch (error) {
console.error("API Error:", error);
}
}
refactorCode();# --- Calling Code (Python) ---
from openai import OpenAI
import os
client = OpenAI(
api_key=os.environ.get("AVALAI_API_KEY"),
base_url="https://api.avalai.ir/v1",
)
# Note the change here: Use indentation for the inner code example
# instead of triple backticks within the prompt string.
prompt = """
Instructions:- Given the React component below, change it so that nonfiction books have red text.
- Return only the refactored React component code in your reply.
- Do not include explanations or markdown code blocks.
- Use four spaces for indentation.
- Keep lines under 80 columns.
Original Code:
const books = [
{ title: 'Dune', category: 'fiction', id: 1 },
{ title: 'Frankenstein', category: 'fiction', id: 2 },
{ title: 'Moneyball', category: 'nonfiction', id: 3 },
];
export default function BookList() {
const listItems = books.map(book =>
<li>
{book.title}
</li>
);
return (
<ul>{listItems}</ul>
);
}
""".strip()
try:
response = client.chat.completions.create(
model="gpt-5.6-sol", # Use a suitable reasoning model from AvalAI
messages=[{"role": "user", "content": prompt}],
temperature=0.1, # Lower temperature for more predictable code output
)
print(response.choices[0].message.content)
except Exception as e:
print(f"API Error: {e}")# --- Calling Code (Bash/cURL) ---
# Note the change here: Use indentation for the inner code example
# instead of triple backticks within the prompt string.
PROMPT=$(
cat <<'EOF'
Instructions:- Given the React component below, change it so that nonfiction books have red text.
- Return only the refactored React component code in your reply.
- Do not include explanations or markdown code blocks.
- Use four spaces for indentation.
- Keep lines under 80 columns.
Original Code:
const books = [
{ title: 'Dune', category: 'fiction', id: 1 },
{ title: 'Frankenstein', category: 'fiction', id: 2 },
{ title: 'Moneyball', category: 'nonfiction', id: 3 },
];
export default function BookList() {
const listItems = books.map(book =>
<li>
{book.title}
</li>
);
return (
<ul>{listItems}</ul>
);
}
EOF
)
# Escape JSON special characters in the prompt
JSON_PROMPT=$(echo "$PROMPT" | jq -Rsa .)
curl https://api.avalai.ir/v1/chat/completions \
-H "Authorization: Bearer $AVALAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-sol",
"messages": [{"role": "user", "content": '"$JSON_PROMPT"'}],
"temperature": 0.1
}'// --- Calling Code (Go) ---
package main
import (
"context"
"fmt"
"os"
"strings"
openai "github.com/openai/openai-go"
)
func main() {
apiKey := os.Getenv("AVALAI_API_KEY")
baseURL := "https://api.avalai.ir/v1"
config := openai.DefaultConfig(apiKey)
config.BaseURL = baseURL
client := openai.NewClientWithConfig(config)
// Note the change here: Use indentation for the inner code example
// instead of triple backticks within the prompt string.
prompt := strings.TrimSpace(`
Instructions:- Given the React component below, change it so that nonfiction books have red text.
- Return only the refactored React component code in your reply.
- Do not include explanations or markdown code blocks.
- Use four spaces for indentation.
- Keep lines under 80 columns.
Original Code:
const books = [
{ title: 'Dune', category: 'fiction', id: 1 },
{ title: 'Frankenstein', category: 'fiction', id: 2 },
{ title: 'Moneyball', category: 'nonfiction', id: 3 },
];
export default function BookList() {
const listItems = books.map(book =>
<li>
{book.title}
</li>
);
return (
<ul>{listItems}</ul>
);
}
`)
temp := float32(0.1)
resp, err := client.CreateChatCompletion(
context.Background(),
openai.ChatCompletionRequest{
Model: "gpt-5.5", // Use a suitable reasoning model from AvalAI
Messages: []openai.ChatCompletionMessage{
{Role: openai.ChatMessageRoleUser, Content: prompt},
},
Temperature: &temp,
},
)
if err != nil {
fmt.Printf("API Error: %v\n", err)
return
}
fmt.Println(resp.Choices[0].Message.Content)
}// --- Calling Code (PHP) ---
<?php
require 'vendor/autoload.php';
use OpenAI\Client;
$apiKey = getenv('AVALAI_API_KEY');
$baseURL = 'https://api.avalai.ir/v1';
// Configure client (example)
$client = OpenAI::client($apiKey);
// Set base URL if needed via factory/config
// Note the change here: Use indentation for the inner code example
// instead of triple backticks within the prompt string.
$prompt = trim(<<<PROMPT
Instructions:- Given the React component below, change it so that nonfiction books have red text.
- Return only the refactored React component code in your reply.
- Do not include explanations or markdown code blocks.
- Use four spaces for indentation.
- Keep lines under 80 columns.
Original Code:
const books = [
{ title: 'Dune', category: 'fiction', id: 1 },
{ title: 'Frankenstein', category: 'fiction', id: 2 },
{ title: 'Moneyball', category: 'nonfiction', id: 3 },
];
export default function BookList() {
const listItems = books.map(book =>
<li>
{book.title}
</li>
);
return (
<ul>{listItems}</ul>
);
}
PROMPT);
try {
$response = $client->chat()->create([
'model' => 'gpt-5.5', // Use a suitable reasoning model from AvalAI
'messages' => [
['role' => 'user', 'content' => $prompt],
],
'temperature' => 0.1,
]);
echo $response->choices[0]->message->content;
} catch (Exception $e) {
echo "API Error: " . $e->getMessage() . "\n";
}
?>Responses API version
Use this version when the selected model supports /v1/responses. It keeps the same refactoring task, moves the application rules into instructions, and uses a low reasoning effort because the edit is constrained and well specified.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AVALAI_API_KEY"],
base_url="https://api.avalai.ir/v1",
)
prompt = """
Given the React component below, change it so that nonfiction books have red text.
Return only the refactored React component code in your reply.
Do not include explanations or markdown code blocks.
Use four spaces for indentation.
Keep lines under 80 columns.
Original Code:
const books = [
{ title: 'Dune', category: 'fiction', id: 1 },
{ title: 'Frankenstein', category: 'fiction', id: 2 },
{ title: 'Moneyball', category: 'nonfiction', id: 3 },
];
export default function BookList() {
const listItems = books.map(book =>
<li>
{book.title}
</li>
);
return (
<ul>{listItems}</ul>
);
}
"""
response = client.responses.create(
model="gpt-5.6-sol",
reasoning={"effort": "low"},
instructions=(
"Formatting re-enabled\n"
"You refactor code precisely. Return only the requested code."
),
input=prompt,
)
print(response.output_text)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.AVALAI_API_KEY,
baseURL: "https://api.avalai.ir/v1",
});
const prompt = `
Given the React component below, change it so that nonfiction books have red text.
Return only the refactored React component code in your reply.
Do not include explanations or markdown code blocks.
Use four spaces for indentation.
Keep lines under 80 columns.
Original Code:
const books = [
{ title: 'Dune', category: 'fiction', id: 1 },
{ title: 'Frankenstein', category: 'fiction', id: 2 },
{ title: 'Moneyball', category: 'nonfiction', id: 3 },
];
export default function BookList() {
const listItems = books.map(book =>
<li>
{book.title}
</li>
);
return (
<ul>{listItems}</ul>
);
}
`;
const response = await client.responses.create({
model: "gpt-5.6-sol",
reasoning: { effort: "low" },
instructions:
"Formatting re-enabled\nYou refactor code precisely. Return only the requested code.",
input: prompt,
});
console.log(response.output_text);PROMPT=$(
cat <<'EOF'
Given the React component below, change it so that nonfiction books have red text.
Return only the refactored React component code in your reply.
Do not include explanations or markdown code blocks.
Use four spaces for indentation.
Keep lines under 80 columns.
Original Code:
const books = [
{ title: 'Dune', category: 'fiction', id: 1 },
{ title: 'Frankenstein', category: 'fiction', id: 2 },
{ title: 'Moneyball', category: 'nonfiction', id: 3 },
];
export default function BookList() {
const listItems = books.map(book =>
<li>
{book.title}
</li>
);
return (
<ul>{listItems}</ul>
);
}
EOF
)
jq -n --arg prompt "$PROMPT" '{
model: "gpt-5.6-sol",
reasoning: {effort: "low"},
instructions: "Formatting re-enabled\nYou refactor code precisely. Return only the requested code.",
input: $prompt
}' | curl https://api.avalai.ir/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AVALAI_API_KEY" \
-d @-messages→input- system message →
instructionsor adeveloperitem choices[0].message.content→response.output_text- for tools and multimodal output, inspect
response.outputby itemtype.
2. Coding (Planning)
Task: Plan and generate code for a simple Python Q&A application.
// --- Calling Code (Node.js) ---
import OpenAI from "openai";
import * as dotenv from "dotenv";
dotenv.config();
const client = new OpenAI({
apiKey: process.env.AVALAI_API_KEY,
baseURL: "https://api.avalai.ir/v1",
});
const prompt = `
I want to build a Python app that takes user questions and looks them up
in a simple key-value store (like a dictionary or JSON file) where they
are mapped to answers. If there is a close match (case-insensitive check),
it retrieves the matched answer. If there isn't, it asks the user to
provide an answer and stores the new question/answer pair.
1. Make a plan for the directory structure (e.g., main script, data file).
2. Return the full Python code for the main script.
3. Return an example JSON structure for the data file.
4. Only supply explanatory text at the very beginning and very end, not mixed within the code or file structure output.
`.trim();
async function planProject() {
try {
const response = await client.chat.completions.create({
model: "deepseek-v4-pro", // Use a suitable reasoning model from AvalAI
messages: [{ role: "user", content: prompt }],
});
console.log(response.choices[0].message.content);
} catch (error) {
console.error("API Error:", error);
}
}
planProject();# --- Calling Code (Python) ---
from openai import OpenAI
import os
client = OpenAI(
api_key=os.environ.get("AVALAI_API_KEY"),
base_url="https://api.avalai.ir/v1",
)
prompt = """
I want to build a Python app that takes user questions and looks them up
in a simple key-value store (like a dictionary or JSON file) where they
are mapped to answers. If there is a close match (case-insensitive check),
it retrieves the matched answer. If there isn't, it asks the user to
provide an answer and stores the new question/answer pair.
1. Make a plan for the directory structure (e.g., main script, data file).
2. Return the full Python code for the main script.
3. Return an example JSON structure for the data file.
4. Only supply explanatory text at the very beginning and very end, not mixed within the code or file structure output.
""".strip()
try:
response = client.chat.completions.create(
model="deepseek-v4-pro", # Use a suitable reasoning model from AvalAI
messages=[{"role": "user", "content": prompt}],
)
print(response.choices[0].message.content)
except Exception as e:
print(f"API Error: {e}")# --- Calling Code (Bash/cURL) ---
PROMPT=$(
cat <<'EOF'
I want to build a Python app that takes user questions and looks them up
in a simple key-value store (like a dictionary or JSON file) where they
are mapped to answers. If there is a close match (case-insensitive check),
it retrieves the matched answer. If there isn't, it asks the user to
provide an answer and stores the new question/answer pair.
1. Make a plan for the directory structure (e.g., main script, data file).
2. Return the full Python code for the main script.
3. Return an example JSON structure for the data file.
4. Only supply explanatory text at the very beginning and very end, not mixed within the code or file structure output.
EOF
)
# Escape JSON special characters
JSON_PROMPT=$(echo "$PROMPT" | jq -Rsa .)
curl https://api.avalai.ir/v1/chat/completions \
-H "Authorization: Bearer $AVALAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-pro",
"messages": [{"role": "user", "content": '"$JSON_PROMPT"'}]
}'// --- Calling Code (Go) ---
package main
import (
"context"
"fmt"
"os"
"strings"
openai "github.com/openai/openai-go"
)
func main() {
apiKey := os.Getenv("AVALAI_API_KEY")
baseURL := "https://api.avalai.ir/v1"
config := openai.DefaultConfig(apiKey)
config.BaseURL = baseURL
client := openai.NewClientWithConfig(config)
prompt := strings.TrimSpace(`
I want to build a Python app that takes user questions and looks them up
in a simple key-value store (like a dictionary or JSON file) where they
are mapped to answers. If there is a close match (case-insensitive check),
it retrieves the matched answer. If there isn't, it asks the user to
provide an answer and stores the new question/answer pair.
1. Make a plan for the directory structure (e.g., main script, data file).
2. Return the full Python code for the main script.
3. Return an example JSON structure for the data file.
4. Only supply explanatory text at the very beginning and very end, not mixed within the code or file structure output.
`)
resp, err := client.CreateChatCompletion(
context.Background(),
openai.ChatCompletionRequest{
Model: "deepseek-v4-pro", // Use a suitable reasoning model from AvalAI
Messages: []openai.ChatCompletionMessage{
{Role: openai.ChatMessageRoleUser, Content: prompt},
},
},
)
if err != nil {
fmt.Printf("API Error: %v\n", err)
return
}
fmt.Println(resp.Choices[0].Message.Content)
}// --- Calling Code (PHP) ---
<?php
require 'vendor/autoload.php';
use OpenAI\Client;
$apiKey = getenv('AVALAI_API_KEY');
$baseURL = 'https://api.avalai.ir/v1';
// Configure client (example)
$client = OpenAI::client($apiKey);
// Set base URL if needed via factory/config
$prompt = trim(<<<PROMPT
I want to build a Python app that takes user questions and looks them up
in a simple key-value store (like a dictionary or JSON file) where they
are mapped to answers. If there is a close match (case-insensitive check),
it retrieves the matched answer. If there isn't, it asks the user to
provide an answer and stores the new question/answer pair.
1. Make a plan for the directory structure (e.g., main script, data file).
2. Return the full Python code for the main script.
3. Return an example JSON structure for the data file.
4. Only supply explanatory text at the very beginning and very end, not mixed within the code or file structure output.
PROMPT);
try {
$response = $client->chat()->create([
'model' => 'deepseek-v4-pro', // Use a suitable reasoning model from AvalAI
'messages' => [
['role' => 'user', 'content' => $prompt],
],
]);
echo $response->choices[0]->message->content;
} catch (Exception $e) {
echo "API Error: " . $e->getMessage() . "\n";
}
?>Responses API version This version uses `gpt-5.5` because `deepseek-v4-pro` may not be enabled for `/v1/responses` in the current AvalAI model data.
Use this version when the selected model supports /v1/responses. It keeps the same planning task and uses medium reasoning effort because the model must design a small structure, write code, and satisfy a strict output contract.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AVALAI_API_KEY"],
base_url="https://api.avalai.ir/v1",
)
prompt = """
I want to build a Python app that takes user questions and looks them up
in a simple key-value store (like a dictionary or JSON file) where they
are mapped to answers. If there is a close match (case-insensitive check),
it retrieves the matched answer. If there isn't, it asks the user to
provide an answer and stores the new question/answer pair.
1. Make a plan for the directory structure (e.g., main script, data file).
2. Return the full Python code for the main script.
3. Return an example JSON structure for the data file.
4. Only supply explanatory text at the very beginning and very end, not mixed within the code or file structure output.
"""
response = client.responses.create(
model="gpt-5.6-sol",
reasoning={"effort": "medium"},
instructions=(
"Formatting re-enabled\n"
"You are a senior Python engineer. Follow the requested output contract exactly."
),
input=prompt,
)
print(response.output_text)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.AVALAI_API_KEY,
baseURL: "https://api.avalai.ir/v1",
});
const prompt = `
I want to build a Python app that takes user questions and looks them up
in a simple key-value store (like a dictionary or JSON file) where they
are mapped to answers. If there is a close match (case-insensitive check),
it retrieves the matched answer. If there isn't, it asks the user to
provide an answer and stores the new question/answer pair.
1. Make a plan for the directory structure (e.g., main script, data file).
2. Return the full Python code for the main script.
3. Return an example JSON structure for the data file.
4. Only supply explanatory text at the very beginning and very end, not mixed within the code or file structure output.
`;
const response = await client.responses.create({
model: "gpt-5.6-sol",
reasoning: { effort: "medium" },
instructions:
"Formatting re-enabled\nYou are a senior Python engineer. Follow the requested output contract exactly.",
input: prompt,
});
console.log(response.output_text);PROMPT=$(
cat <<'EOF'
I want to build a Python app that takes user questions and looks them up
in a simple key-value store (like a dictionary or JSON file) where they
are mapped to answers. If there is a close match (case-insensitive check),
it retrieves the matched answer. If there isn't, it asks the user to
provide an answer and stores the new question/answer pair.
1. Make a plan for the directory structure (e.g., main script, data file).
2. Return the full Python code for the main script.
3. Return an example JSON structure for the data file.
4. Only supply explanatory text at the very beginning and very end, not mixed within the code or file structure output.
EOF
)
jq -n --arg prompt "$PROMPT" '{
model: "gpt-5.6-sol",
reasoning: {effort: "medium"},
instructions: "Formatting re-enabled\nYou are a senior Python engineer. Follow the requested output contract exactly.",
input: $prompt
}' | curl https://api.avalai.ir/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AVALAI_API_KEY" \
-d @-messages→input- system message →
instructionsor adeveloperitem choices[0].message.content→response.output_text- for tools and multimodal output, inspect
response.outputby itemtype.
3. STEM Research
Task: Ask for potential compounds for antibiotic research.
// --- Calling Code (Node.js) ---
import OpenAI from "openai";
import * as dotenv from "dotenv";
dotenv.config();
const client = new OpenAI({
apiKey: process.env.AVALAI_API_KEY,
baseURL: "https://api.avalai.ir/v1",
});
const prompt = `
What are three compounds or classes of compounds we should consider
investigating further to advance research into new antibiotics,
especially against resistant bacteria? Briefly explain why each
is promising.
`.trim();
async function researchQuery() {
try {
const response = await client.chat.completions.create({
model: "gemini-3.1-pro-preview", // Use a suitable reasoning model from AvalAI
messages: [{ role: "user", content: prompt }],
});
console.log(response.choices[0].message.content);
} catch (error) {
console.error("API Error:", error);
}
}
researchQuery();# --- Calling Code (Python) ---
from openai import OpenAI
import os
client = OpenAI(
api_key=os.environ.get("AVALAI_API_KEY"),
base_url="https://api.avalai.ir/v1",
)
prompt = """
What are three compounds or classes of compounds we should consider
investigating further to advance research into new antibiotics,
especially against resistant bacteria? Briefly explain why each
is promising.
""".strip()
try:
response = client.chat.completions.create(
model="gemini-3.1-pro-preview", # Use a suitable reasoning model from AvalAI
messages=[{"role": "user", "content": prompt}],
)
print(response.choices[0].message.content)
except Exception as e:
print(f"API Error: {e}")# --- Calling Code (Bash/cURL) ---
PROMPT=$(
cat <<'EOF'
What are three compounds or classes of compounds we should consider
investigating further to advance research into new antibiotics,
especially against resistant bacteria? Briefly explain why each
is promising.
EOF
)
# Escape JSON special characters
JSON_PROMPT=$(echo "$PROMPT" | jq -Rsa .)
curl https://api.avalai.ir/v1/chat/completions \
-H "Authorization: Bearer $AVALAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.1-pro-preview",
"messages": [{"role": "user", "content": '"$JSON_PROMPT"'}]
}'// --- Calling Code (Go) ---
package main
import (
"context"
"fmt"
"os"
"strings"
openai "github.com/openai/openai-go"
)
func main() {
apiKey := os.Getenv("AVALAI_API_KEY")
baseURL := "https://api.avalai.ir/v1"
config := openai.DefaultConfig(apiKey)
config.BaseURL = baseURL
client := openai.NewClientWithConfig(config)
prompt := strings.TrimSpace(`
What are three compounds or classes of compounds we should consider
investigating further to advance research into new antibiotics,
especially against resistant bacteria? Briefly explain why each
is promising.
`)
resp, err := client.CreateChatCompletion(
context.Background(),
openai.ChatCompletionRequest{
Model: "gemini-3.1-pro-preview", // Use a suitable reasoning model from AvalAI
Messages: []openai.ChatCompletionMessage{
{Role: openai.ChatMessageRoleUser, Content: prompt},
},
},
)
if err != nil {
fmt.Printf("API Error: %v\n", err)
return
}
fmt.Println(resp.Choices[0].Message.Content)
}// --- Calling Code (PHP) ---
<?php
require 'vendor/autoload.php';
use OpenAI\Client;
$apiKey = getenv('AVALAI_API_KEY');
$baseURL = 'https://api.avalai.ir/v1';
// Configure client (example)
$client = OpenAI::client($apiKey);
// Set base URL if needed via factory/config
$prompt = trim(<<<PROMPT
What are three compounds or classes of compounds we should consider
investigating further to advance research into new antibiotics,
especially against resistant bacteria? Briefly explain why each
is promising.
PROMPT);
try {
$response = $client->chat()->create([
'model' => 'gemini-3.1-pro-preview', // Use a suitable reasoning model from AvalAI
'messages' => [
['role' => 'user', 'content' => $prompt],
],
]);
echo $response->choices[0]->message->content;
} catch (Exception $e) {
echo "API Error: " . $e->getMessage() . "\n";
}
?>Responses API version This version uses `gpt-5.5` because `gemini-3.1-pro-preview` may not be enabled for `/v1/responses` in the current AvalAI model data.
Use this version when the selected model supports /v1/responses. It keeps the same research-style prompt, but adds an explicit safety boundary so the answer stays high level and does not provide synthesis or wet-lab instructions.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AVALAI_API_KEY"],
base_url="https://api.avalai.ir/v1",
)
prompt = """
What are three compounds or classes of compounds we should consider
investigating further to advance research into new antibiotics,
especially against resistant bacteria? Briefly explain why each is promising.
"""
response = client.responses.create(
model="gpt-5.6-sol",
reasoning={"effort": "high"},
instructions=(
"Formatting re-enabled\n"
"Answer at a high scientific level. Do not include synthesis steps, "
"dosages, protocols, or operational wet-lab instructions."
),
input=prompt,
)
print(response.output_text)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.AVALAI_API_KEY,
baseURL: "https://api.avalai.ir/v1",
});
const prompt = `
What are three compounds or classes of compounds we should consider
investigating further to advance research into new antibiotics,
especially against resistant bacteria? Briefly explain why each is promising.
`;
const response = await client.responses.create({
model: "gpt-5.6-sol",
reasoning: { effort: "high" },
instructions:
"Formatting re-enabled\nAnswer at a high scientific level. Do not include synthesis steps, dosages, protocols, or operational wet-lab instructions.",
input: prompt,
});
console.log(response.output_text);PROMPT=$(
cat <<'EOF'
What are three compounds or classes of compounds we should consider
investigating further to advance research into new antibiotics,
especially against resistant bacteria? Briefly explain why each is promising.
EOF
)
jq -n --arg prompt "$PROMPT" '{
model: "gpt-5.6-sol",
reasoning: {effort: "high"},
instructions: "Formatting re-enabled\nAnswer at a high scientific level. Do not include synthesis steps, dosages, protocols, or operational wet-lab instructions.",
input: $prompt
}' | curl https://api.avalai.ir/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AVALAI_API_KEY" \
-d @-messages→input- system message →
instructionsor adeveloperitem choices[0].message.content→response.output_text- for tools and multimodal output, inspect
response.outputby itemtype.
Use Case Examples
Explore the AvalAI Cookbook (if available) or community resources for more examples of applying reasoning models to tasks like data validation, routine generation, and complex analysis.