Developer Dashboard

Reasoning Models

Explore advanced reasoning and problem-solving models available through AvalAI.

Introduction

Reasoning models are large language models trained to spend more compute on complex reasoning before they answer. Treat their reasoning as an internal process: ask for final answers, concise rationales, citations, or checklists, but do not ask the model to reveal hidden chain-of-thought. They excel in complex problem solving, coding, scientific reasoning, and multi-step planning for agentic workflows.

AvalAI offers access to several reasoning-capable models from different providers:

OpenAI Models:

  • gpt-5.6-sol: OpenAI's newest highest-capability GPT-5.6 flagship for hard agentic coding, knowledge work, scientific reasoning, and tool orchestration, 1M context window
  • gpt-5.6-terra: Balanced GPT-5.6 model for everyday production reasoning and agentic workflows, 1M context window
  • gpt-5.6-luna: Cost-efficient GPT-5.6 model for high-volume reasoning, support, and document workflows, 1M context window
  • gpt-5.5: Previous OpenAI flagship model with state-of-the-art reasoning across agentic coding, knowledge work, computer use, and scientific research (configurable effort: none/low/medium/high/xhigh), 1M context window
  • gpt-5.4-pro: OpenAI's highest-reasoning GPT-5.4 model for complex professional work, 1.05M context window, effort levels medium/high/xhigh
  • gpt-5.4: Frontier reasoning and agentic workflow model with configurable effort (none/low/medium/high/xhigh), 1.05M context window
  • gpt-5.4-mini: Fast, cost-efficient model with reasoning support (none, low, medium effort levels), 400K context
  • gpt-5.4-nano: Fastest and most affordable model with basic reasoning (none, low effort levels), 400K context
  • gpt-5-pro: OpenAI's advanced reasoning model with extended thinking capabilities for expert-level problem solving (Tier 2+ users, Responses API only)
  • gpt-5.3-codex: OpenAI's most capable agentic coding model with reasoning tokens support (Responses API only)
  • o1-pro: The most capable traditional reasoning model, but also the most expensive
  • o4-mini: Enhanced reasoning with improved efficiency
  • o3: Balanced reasoning capabilities with good performance
  • o3-mini: A smaller, faster model, generally less expensive per token

Google's Gemini Models:

  • gemini-3.6-flash: Google's July 2026 workhorse Flash reasoning model for agentic coding, knowledge work, multimodal analysis, efficient tool use, and 1M context
  • gemini-3.5-flash-lite: Google's fastest and most cost-effective Gemini 3.5-class reasoning model for high-throughput subagents, document parsing, extraction, and low-latency workloads
  • gemini-3.5-flash: Google's May 2026 Flash reasoning model with configurable thinking levels, strong coding and agentic tool-use performance, multimodal input, and 1M context
  • gemini-3.1-pro-preview: Google's advanced Pro-class model (Feb 2026) with natively multimodal reasoning, strong agentic performance, advanced coding, and long-context understanding
  • gemini-3.1-flash-lite: Stable cost-efficient reasoning model optimized for high-frequency, lightweight agentic tasks with extremely low latency
  • gemini-3.1-flash-lite-preview: Preview alias for Gemini 3.1 Flash-Lite with the same pricing and capabilities
  • gemini-2.5-pro: Features enhanced thinking and reasoning capabilities
  • gemini-2.5-flash: Google's first hybrid reasoning model with configurable thinking budgets

Anthropic Models:

  • claude-opus-5: Anthropic's new flagship Opus model with a 1M-token input window, 128K output capacity, adaptive thinking, configurable effort, stronger verification, and improved long-horizon coding, knowledge work, computer use, and scientific analysis
  • claude-opus-4-8: Previous Opus flagship with adaptive thinking, 5 effort levels (low/medium/high/xhigh/max), default high effort, mid-conversation system messages, and strong long-horizon agentic coding
  • claude-opus-4-7: Previous flagship with adaptive thinking, 5 effort levels (low/medium/high/xhigh/max), and task budgets
  • claude-sonnet-5: Anthropic's most agentic Sonnet model with adaptive thinking and configurable effort levels; higher effort can match Opus 4.8 on some tasks at lower Sonnet-tier prices
  • claude-sonnet-4-6: Strong reasoning capabilities with better efficiency
  • claude-haiku-4-5: Fast, efficient reasoning for cost-sensitive deployments

Moonshot AI Models:

  • kimi-k3: Moonshot AI's new 2.8T-parameter flagship with a 1M-token context window, native vision, always-on reasoning, long-horizon coding, structured output, and tool calling
  • kimi-latest: Stable alias that now resolves to kimi-k3 with the same capabilities and pricing

For Kimi K3, use the top-level reasoning_effort: "max" field. K3 currently supports only the max effort level and keeps thinking enabled. Do not reuse the older K2.x thinking parameter or send fixed sampling fields such as temperature and top_p. In multi-turn and tool workflows, append the complete assistant message so the reasoning and tool-call context remains intact.

DeepSeek Models:

  • deepseek-v4-pro: DeepSeek's flagship reasoning model (1.6T total / 49B active params) with open-source SOTA Agentic Coding, 1M context window, provider-exposed reasoning_content, reasoning_effort: "high"/"max"
  • deepseek-v4-flash: Stable ID now automatically backed by the official DeepSeek-V4-Flash-0731 release (284B total / 13B active params), with a 1M context window, stronger long-horizon coding and tool use, and reasoning_effort: "low"/"high"/"max"; no code or pricing change is required
  • deepseek-reasoner: Legacy alias, now routes to deepseek-v4-pro (retiring July 24, 2026)
  • deepseek-chat: Legacy alias, now routes to deepseek-v4-flash (retiring July 24, 2026)

XAI Models:

  • grok-4.5: XAI's new flagship model for coding, agentic tasks, and knowledge work with 1M context, fast serving, Chat Completions support, and partial Responses support
  • grok-4.3: XAI's flagship reasoning model with 1M context window, function calling, structured outputs, and context-aware pricing above 200K tokens
  • grok-4.20-reasoning: Stable release with industry-leading speed and built-in reasoning, 2M context window, lowest hallucination rate
  • grok-4.20-non-reasoning: Stable non-reasoning variant for tasks not requiring extended internal reasoning, 2M context window

MiniMax Models:

  • minimax-m3: New flagship model with frontier coding and agentic capabilities, 1M context window (MSA architecture), native multimodal input, and toggleable thinking
  • minimax-m2.7: Revolutionary self-evolution reasoning model (first model to deeply participate in its own evolution), 56.22% SWE-Pro, Agent Teams support
  • minimax-m2.7-highspeed: Ultra-fast self-evolution variant (~100 tokens per second output speed)
  • minimax-m2.5: Earlier flagship model with SOTA coding (80.2% SWE-Bench Verified) and real-world productivity
  • minimax-m2.5-lightning: Ultra-fast reasoning variant (~100 tokens per second output speed)
  • minimax-m2.1: Flagship reasoning model with o3-level performance and 20x efficiency
  • minimax-m2.1-lightning: Fast reasoning variant (~100 tokens per second output speed)
  • minimax-m2: Versatile reasoning model with strong general capabilities

Z.AI Models:

  • glm-5.2: Latest flagship reasoning model with 1M context, frontier coding, and long-horizon agentic engineering
  • glm-5.1: State-of-the-art reasoning with 58.4% SWE-Bench Pro, long-horizon optimization for agentic engineering
  • glm-5v-turbo: Multimodal reasoning and vision understanding with high-throughput visual processing
  • glm-5-turbo: OpenClaw-optimized reasoning for tool invocation, persistent tasks, and long-chain execution
  • glm-5: Flagship model with SOTA coding and agentic engineering capabilities

Alibaba Models:

  • qwen3.8-max: Newest 2.4T-parameter flagship for long-horizon coding, professional work, multimodal reasoning, and agentic planning, with hybrid thinking through enable_thinking, a 1M context window, and up to 128K output
  • qwen3.7-max: Previous flagship agent-foundation model, hybrid thinking via enable_thinking, 92.4 GPQA Diamond, 97.1 HMMT, 1M context
  • qwen3-max: Flagship Qwen3 Max model for complex reasoning and agentic workflows, hybrid thinking via enable_thinking, 262K context
  • qwen3.6-plus: Agentic coding model with 78.8% SWE-bench Verified, 1M context default, frontier web development
  • qwen3.6-flash: Fast hybrid-thinking model with enable_thinking toggle, 1M context window
  • qwen3.6-max-preview: Flagship preview with hybrid thinking, most capable Qwen3.6 model
  • qwen3.6-35b-a3b: Open-weight MoE (35B total/3B active) with thinking mode, 256K context
  • qwen3.6-27b: Dense vision-language reasoning model with 256K context
  • qwen3-coder-next: 80B MoE model (10B active) optimized for coding agents with 1M context

Fireworks.ai Models:

  • nemotron-3-ultra: NVIDIA's flagship large-scale Nemotron model for complex reasoning and agentic workflows, served via Fireworks.ai

Note: Some advanced models like o1-pro might have unique features and specific API endpoints (e.g., AvalAI's equivalent of the Responses API). Refer to the specific model documentation and the AvalAI API Reference for details.

When to use reasoning models

Reasoning models are best for tasks where correctness depends on planning, ambiguity resolution, or careful trade-offs. Use them when you need:

  • Complex problem solving: math, science, finance, legal, policy, or engineering decisions with many constraints.
  • Long-context synthesis: finding relationships across contracts, reports, transcripts, tickets, or retrieved documents.
  • Agentic planning: deciding which tools to call, decomposing a workflow, or assigning steps to faster execution models.
  • Code review and debugging: inspecting multi-file diffs, root-causing failures, or validating generated patches.
  • Evaluation: judging model answers against nuanced rubrics or gold-standard requirements.

Think of reasoning models as planners and faster GPT-style models as workhorses. Use the planner for ambiguity, tool strategy, policy interpretation, long-context synthesis, or final validation. Use the workhorse for extraction, rewriting, classification, formatting, and other well-defined execution steps where speed and cost matter more than deep reasoning. Many production systems pair both: a reasoning model plans or validates, while a lower-latency model executes simple steps.

Prompting reasoning models effectively

Reasoning models need a different prompt style from classic GPT-style models:

  • Keep the prompt simple and direct; start zero-shot before adding examples.
  • Use a developer message for application rules and a user message for the task.
  • Use Markdown headings, XML tags, or other delimiters to separate rules, context, and examples.
  • Be specific about constraints, success criteria, tools available, and what the final answer should contain.
  • Avoid asking for hidden chain-of-thought or "think step by step"; ask for a short rationale, answer checklist, or cited evidence instead.
  • If a reasoning-model snapshot suppresses Markdown but your UI needs Markdown, begin the developer message with Formatting re-enabled.

Provider features vary. OpenAI models use reasoning.effort on /v1/responses; Kimi K3 uses top-level reasoning_effort: "max"; DeepSeek-V4-Flash-0731 accepts top-level reasoning_effort values low, high, and max; Alibaba Qwen models such as qwen3.8-max use enable_thinking with route-specific streaming requirements. Check the provider section for the selected model before copying parameters across families.

Get Started with Reasoning

For OpenAI reasoning models available through AvalAI, use /v1/responses when the model supports it. Responses preserves typed output items, supports reasoning.effort, and can keep relevant reasoning items across tool calls with previous_response_id or by replaying prior output items. Keep Chat Completions examples for provider-specific routes or legacy integrations that do not expose Responses yet.

Responses-first reasoning checklist

  • Start with gpt-5.6-terra or gpt-5.6-sol and reasoning: {"effort": "medium"} for balanced quality, latency, and cost; try low for faster support or drafting flows and high/xhigh only when evals justify the added tokens.
  • Reserve enough max_output_tokens for both visible output and hidden reasoning tokens. If a response returns status: "incomplete" with incomplete_details.reason: "max_output_tokens", the model may have spent the entire budget on reasoning and produced no visible text. Increase the budget, lower reasoning.effort to low or none when supported, or simplify the task.
  • When a reasoning model calls tools, set store: true when your retention policy allows it, then pass back the reasoning items and function-call items from the prior response, or continue with previous_response_id.
  • For stateless or zero-retention style flows, include reasoning.encrypted_content when supported so encrypted reasoning items can be round-tripped.
  • Request reasoning.summary only when the selected model and account support summaries; summaries are for observability, not raw chain-of-thought.
  • Prompt reasoning models with goals, constraints, available tools, and success criteria. Avoid asking for hidden chain-of-thought; ask for a concise final explanation or checklist instead.

Latency and visible progress

For hard tasks, raising reasoning.effort can improve answer quality but also delays the first useful token. OpenAI's reasoning guidance recommends asking the model for a short preamble when you need faster visible progress. In AvalAI, use this pattern for streamed developer tools, review assistants, and long-running agentic workflows:

  • Ask for one short progress sentence before deeper analysis, not hidden chain-of-thought.
  • Render the preamble as status text; do not treat it as the final answer.
  • Keep the final answer contract explicit: decision, evidence, checklist, patch plan, or next action.
  • If the route supports phase, preserve phase: "commentary" preambles and phase: "final_answer" final responses when replaying state.
python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AVALAI_API_KEY"],
    base_url="https://api.avalai.ir/v1",
)

response = client.responses.create(
    model="gpt-5.6-sol",
    reasoning={"effort": "high"},
    instructions=(
        "For difficult work, first give one short progress sentence. "
        "Do not reveal hidden chain-of-thought. Finish with a concise "
        "decision, evidence, and verification checklist."
    ),
    input="Review this migration plan for production risks: ...",
)

print(response.output_text)
javascript
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.AVALAI_API_KEY,
  baseURL: "https://api.avalai.ir/v1",
});

const response = await client.responses.create({
  model: "gpt-5.6-sol",
  reasoning: { effort: "high" },
  instructions:
    "For difficult work, first give one short progress sentence. Do not reveal hidden chain-of-thought. Finish with a concise decision, evidence, and verification checklist.",
  input: "Review this migration plan for production risks: ...",
});

console.log(response.output_text);

Long-running tool flows and phase

OpenAI's current Responses guidance recommends preserving the assistant phase field for long-running or tool-heavy GPT-5.5/GPT-5.4 workflows. Treat this as model- and route-dependent in AvalAI: prefer previous_response_id when storage is allowed, and if you manually replay assistant history, keep any original phase values unchanged.

Use phase: "commentary" for intermediate assistant updates before tool calls, and phase: "final_answer" for the completed answer. Do not add phase to user messages. Dropping phase metadata can make an intermediate note look like the final answer in multi-step flows.

python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AVALAI_API_KEY"],
    base_url="https://api.avalai.ir/v1",
)

response = client.responses.create(
    model="gpt-5.6-sol",
    input=[
        {
            "role": "assistant",
            "phase": "commentary",
            "content": "I will inspect the logs before recommending a fix.",
        },
        {
            "role": "assistant",
            "phase": "final_answer",
            "content": "Root cause: cache invalidation race.",
        },
        {"role": "user", "content": "Now give me a rollout-safe fix plan."},
    ],
)

print(response.output_text)
javascript
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.AVALAI_API_KEY,
  baseURL: "https://api.avalai.ir/v1",
});

const response = await client.responses.create({
  model: "gpt-5.6-sol",
  input: [
    {
      role: "assistant",
      phase: "commentary",
      content: "I will inspect the logs before recommending a fix.",
    },
    {
      role: "assistant",
      phase: "final_answer",
      content: "Root cause: cache invalidation race.",
    },
    { role: "user", content: "Now give me a rollout-safe fix plan." },
  ],
});

console.log(response.output_text);

Example: Using a reasoning model

bash
PROMPT='Write a bash script that takes a matrix represented as a string with format '\''[1,2],[3,4],[5,6]'\'' and prints the transpose in the same format.'

curl https://api.avalai.ir/v1/responses \
  -H "Authorization: Bearer $AVALAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
 "model": "gpt-5.6-sol",
 "reasoning": {"effort": "medium"},
 "input": [
 {
 "role": "user",
 "content": "'"$PROMPT"'"
 }
 ]
 }'
python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AVALAI_API_KEY"],
    base_url="https://api.avalai.ir/v1",
)

prompt = """
Write a bash script that takes a matrix represented as a string with
format '[1,2],[3,4],[5,6]' and prints the transpose in the same format.
"""

try:
    response = client.responses.create(
        model="gpt-5.6-sol",
        reasoning={"effort": "medium"},
        input=[{"role": "user", "content": prompt}],
    )
    print(response.output_text)

except Exception as e:
    print(f"An API error occurred: {e}")
javascript
import OpenAI from "openai"; // Use the standard OpenAI library configured for AvalAI
import * as dotenv from "dotenv";
dotenv.config();

const client = new OpenAI({
  apiKey: process.env.AVALAI_API_KEY,
  baseURL: "https://api.avalai.ir/v1", // AvalAI endpoint
});

const prompt = `
Write a bash script that takes a matrix represented as a string with
format '[1,2],[3,4],[5,6]' and prints the transpose in the same format.
`;

async function runReasoning() {
  try {
    const response = await client.responses.create({
      model: "gpt-5.6-sol",
      reasoning: { effort: "medium" },
      input: [{ role: "user", content: prompt }],
    });
    console.log(response.output_text);
  } catch (error) {
    console.error("An API error occurred:", error);
  }
}

runReasoning();
go
package main

import (
	"bytes"
	"encoding/json"
	"fmt"
	"io"
	"net/http"
	"os"
)

func main() {
	prompt := `
Write a bash script that takes a matrix represented as a string with
format '[1,2],[3,4],[5,6]' and prints the transpose in the same format.
`

	payload := map[string]any{
		"model":     "gpt-5.5",
		"reasoning": map[string]string{"effort": "medium"},
		"input": []map[string]string{
			{
				"role":    "user",
				"content": prompt,
			},
		},
	}

	body, err := json.Marshal(payload)
	if err != nil {
		panic(err)
	}

	req, err := http.NewRequest("POST", "https://api.avalai.ir/v1/responses", bytes.NewBuffer(body))
	if err != nil {
		panic(err)
	}
	req.Header.Set("Authorization", "Bearer "+os.Getenv("AVALAI_API_KEY"))
	req.Header.Set("Content-Type", "application/json")

	resp, err := http.DefaultClient.Do(req)

	if err != nil {
		fmt.Printf("Responses API error: %v\n", err)
		return
	}
	defer resp.Body.Close()

	responseBody, err := io.ReadAll(resp.Body)
	if err != nil {
		panic(err)
	}

	fmt.Println(string(responseBody))
}
php
<?php
$apiKey = getenv('AVALAI_API_KEY');
$prompt = <<<PROMPT
Write a bash script that takes a matrix represented as a string with
format '[1,2],[3,4],[5,6]' and prints the transpose in the same format.
PROMPT;

$payload = [
    'model' => 'gpt-5.5',
    'reasoning' => ['effort' => 'medium'],
    'input' => [
        ['role' => 'user', 'content' => $prompt],
    ],
];

$ch = curl_init('https://api.avalai.ir/v1/responses');
curl_setopt_array($ch, [
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_POST => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $apiKey,
        'Content-Type: application/json',
    ],
    CURLOPT_POSTFIELDS => json_encode($payload),
]);

$response = curl_exec($ch);
curl_close($ch);

echo $response;

?>

Reasoning Effort

Some reasoning models accept parameters like reasoning.effort to guide the amount of internal reasoning performed before generating a response. Potential values include:

  • none: Disables explicit reasoning where supported, favoring speed.
  • minimal: Uses the smallest supported reasoning budget when a model exposes this value.
  • low: Favors speed and lower token usage.
  • medium (often default): Balances speed and reasoning quality.
  • high: Favors more thorough reasoning, potentially using more tokens and taking longer.
  • xhigh or max: Uses the most reasoning for the hardest coding, math, and agentic tasks where supported.

Consult the specific model documentation on AvalAI for supported parameters and their effects.

Choosing effort in production

Treat effort as a latency/cost/quality tuning knob, not as the first fix for a weak prompt. Start with the lowest setting that passes your evals, then raise it only for tasks where the extra reasoning tokens produce measurable quality gains.

WorkloadSuggested starting pointWhy
Voice, classification, simple retrievalnone or low when supportedPrioritizes first-token latency and lower token use.
Customer support, drafting, tool planninglow or mediumLeaves room for tool strategy without over-spending.
Coding, research, spreadsheet/document analysismediumMatches the balanced default for current OpenAI reasoning guidance such as gpt-5.5.
Deep research, security review, hard debugginghigh or xhigh after evalsUseful when accuracy matters more than latency and cost.

Defaults are model-dependent. Do not assume that medium is universal across OpenAI, Anthropic, Gemini, DeepSeek, XAI, or other providers; record the chosen effort in your eval traces and compare quality, latency, output_tokens, and output_tokens_details.reasoning_tokens before changing production defaults.

How Reasoning Works

Reasoning models utilize reasoning tokens internally, in addition to the standard input and output tokens. These tokens represent the model's "thought process" – breaking down the prompt, exploring approaches, and planning the response. After this internal reasoning phase, the model generates the final visible output tokens and typically discards the reasoning tokens from its context memory for subsequent turns.

Diagram showing reasoning tokens used internally but not kept in context(Diagram Source: OpenAI)

Important: Reasoning tokens can be hidden from the final response text and still be billable and occupy the model's context window during generation. Across providers and models, they are billed at the selected model's output-token rate. When output_tokens already includes them, output_tokens_details.reasoning_tokens is a breakdown, so adding reasoning_tokens to output_tokens would double-count them. If a route reports visible output and reasoning separately, apply the same output-token rate to both counts.

Provider-Specific Reasoning Settings

Different model providers implement reasoning capabilities in their own ways. Here's how to use reasoning features with each provider through AvalAI:

Gemini Models Reasoning Settings

Google's current Gemini 3.6, 3.5, and 3.1 reasoning-capable models support configurable thinking through Gemini generationConfig.thinkingConfig. For gemini-3.6-flash, gemini-3.5-flash-lite, gemini-3.5-flash, gemini-3.1-pro-preview, gemini-3.1-flash-lite, and gemini-3.1-flash-lite-preview, use thinkingLevel to balance quality, latency, and cost.

When using OpenAI client libraries with AvalAI, pass Gemini-specific settings through the extra_body parameter:

python
# Python example - Gemini thinking levels
response = client.chat.completions.create(
    model="gemini-3.6-flash",
    messages=[
        {
            "role": "user",
            "content": "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ...",
        }
    ],
    extra_body={"generationConfig": {"thinkingConfig": {"thinkingLevel": "high"}}},
)
javascript
// JavaScript example - Gemini thinking levels
const response = await client.chat.completions.create({
  model: "gemini-3.6-flash",
  messages: [
    {
      role: "user",
      content: "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ...",
    },
  ],
  // @ts-expect-error extra_body is supported by AvalAI for Gemini-specific options
  extra_body: {
    generationConfig: {
      thinkingConfig: { thinkingLevel: "high" },
    },
  },
});
Responses API version

Gemini 3.6 Flash and Gemini 3.5 Flash-Lite have partial /v1/responses support. Use this version when the parameters and tools required by your workflow are supported; messages moves to input, and the final text is read from response.output_text.

python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AVALAI_API_KEY"],
    base_url="https://api.avalai.ir/v1",
)

response = client.responses.create(
    model="gemini-3.6-flash",
    instructions="You are a helpful assistant.",
    input="Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ...",
)

print(response.output_text)
javascript
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.AVALAI_API_KEY,
  baseURL: "https://api.avalai.ir/v1",
});

const response = await client.responses.create({
  model: "gemini-3.6-flash",
  instructions: "You are a helpful assistant.",
  input: "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ...",
});

console.log(response.output_text);
bash
curl https://api.avalai.ir/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AVALAI_API_KEY" \
  -d '
  {
    "model": "gemini-3.6-flash",
    "input": "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ...",
    "instructions": "You are a helpful assistant."
  }'
  • messagesinput
  • system message → instructions or a developer item
  • choices[0].message.contentresponse.output_text
  • for tools and multimodal output, inspect response.output by item type.

The Gemini 3.6/3.5/3.1 thinking settings include:

  • thinkingLevel: Controls reasoning depth. Supported levels depend on the selected model and current route.
  • Higher thinking levels generally improve multi-step reasoning and tool-use quality, but can increase latency and token usage.
  • Use a mid-range supported level for balanced production defaults and a higher supported level for complex coding, agentic, or analytical tasks.

For the native Gemini v1beta endpoint, pass the same configuration directly in the request body:

bash
curl https://api.avalai.ir/v1beta/models/gemini-3.6-flash:generateContent \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AVALAI_API_KEY" \
  -d '{
    "contents": [{"role": "user", "parts": [{"text": "Plan a resilient migration from a monolith to microservices."}]}],
    "generationConfig": {
      "thinkingConfig": {"thinkingLevel": "high"}
    }
  }'

Legacy note

thinking_budget is specific to Gemini 2.5 Flash. For the most current information, reference the official Google AI documentation.

Gemini 2.5 Flash uses budget-based thinking controls:

python
# Python example - Gemini 2.5 Flash thinking budget
response = client.chat.completions.create(
    model="gemini-2.5-flash",
    messages=[
        {
            "role": "user",
            "content": "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ...",
        }
    ],
    extra_body={"thinking": {"type": "enabled", "budget_tokens": 2000}},
)
javascript
// JavaScript example - Gemini 2.5 Flash thinking budget
const responseAlt = await client.chat.completions.create({
  model: "gemini-2.5-flash",
  messages: [{ role: "user", content: "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ..." }],
  // @ts-expect-error thinking is an undocumented provider-specific parameter
  thinking: { type: "enabled", budget_tokens: 2000 },
});
Responses API version This version uses `gpt-5.5` because `gemini-2.5-flash` may not be enabled for `/v1/responses` in the current AvalAI model data.

Use this version when the selected model supports /v1/responses. messages moves to input, and the final text is read from response.output_text.

python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AVALAI_API_KEY"],
    base_url="https://api.avalai.ir/v1",
)

response = client.responses.create(
    model="gpt-5.6-sol",
    input=[
        {
            "role": "user",
            "content": [
                {
                    "type": "input_text",
                    "text": "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ...",
                },
                {"type": "input_file", "file_id": "file_abc123"},
            ],
        }
    ],
)

print(response.output_text)
javascript
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.AVALAI_API_KEY,
  baseURL: "https://api.avalai.ir/v1",
});

const response = await client.responses.create({
  model: "gpt-5.6-sol",
  input: [
    {
      role: "user",
      content: [
        { type: "input_text", text: "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ..." },
        { type: "input_file", file_id: "file_abc123" },
      ],
    },
  ],
});

console.log(response.output_text);
bash
curl https://api.avalai.ir/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AVALAI_API_KEY" \
  -d '
  {
    "model": "gpt-5.6-sol",
    "input": [
      {
        "role": "user",
        "content": [
          {
            "type": "input_text",
            "text": "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ..."
          },
          {
            "type": "input_file",
            "file_id": "file_abc123"
          }
        ]
      }
    ]
  }'
  • messagesinput
  • system message → instructions or a developer item
  • choices[0].message.contentresponse.output_text
  • for tools and multimodal output, inspect response.output by item type.

By setting appropriate thinking levels or budgets, you can control both the depth of reasoning and the cost of your API calls.

Anthropic Models Reasoning Settings

Anthropic's Claude models (specifically claude-opus-5, claude-sonnet-5, claude-opus-4-8, claude-opus-4-7, claude-sonnet-4-6, claude-haiku-4-5) support configurable reasoning through provider-specific thinking settings. For Claude Opus 5, use thinking: {"type": "adaptive"} together with output_config.effort; do not send a fixed extended-thinking budget. Start at medium or high, then increase to xhigh or max only when evaluations show a material improvement in task success.

python
response = client.chat.completions.create(
    model="claude-opus-5",
    messages=[
        {
            "role": "user",
            "content": "Analyze this migration design, identify failure modes, and verify the rollback plan.",
        }
    ],
    extra_body={
        "thinking": {"type": "adaptive"},
        "output_config": {"effort": "high"},
    },
)
Responses API version Claude Opus 5 has partial `/v1/responses` support; verify required fields and tools before production use.

Use this version when the selected model supports /v1/responses. messages moves to input, and the final text is read from response.output_text.

python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AVALAI_API_KEY"],
    base_url="https://api.avalai.ir/v1",
)

response = client.responses.create(
    model="gpt-5.6-sol",
    instructions="You are a helpful assistant.",
    input="Analyze this migration design, identify failure modes, and verify the rollback plan.",
)

print(response.output_text)
  • messagesinput
  • system message → instructions or a developer item
  • choices[0].message.contentresponse.output_text
  • for tools and multimodal output, inspect response.output by item type.

Anthropic effort controls are provider-specific; do not copy Gemini thinking parameters or OpenAI reasoning.effort fields without checking the selected route. Request concise rationales or verification evidence rather than hidden chain-of-thought.

OpenAI Models Reasoning Settings

OpenAI's latest reasoning-capable models (gpt-5.5, gpt-5.4-pro, gpt-5.4, gpt-5.4-mini, gpt-5.4-nano, gpt-5.3-codex, gpt-5-pro, o4-mini, o3, and o3-mini) have built-in reasoning capabilities that are automatically engaged when needed. For these models, you may see reasoning tokens included in your usage statistics, but the reasoning process is more integrated into the model's operation.

For more advanced control, some endpoints may support parameters like reasoning.effort to guide the amount of internal reasoning:

python
response = client.chat.completions.create(
    model="gpt-5.6-sol",
    messages=[
        {
            "role": "user",
            "content": "Design an algorithm to solve this optimization problem: ...",
        }
    ],
    extra_body={"reasoning": {"effort": "high"}},  # Request more thorough reasoning
)
Responses API version

Use this version when the selected model supports /v1/responses. messages moves to input, and the final text is read from response.output_text.

python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AVALAI_API_KEY"],
    base_url="https://api.avalai.ir/v1",
)

response = client.responses.create(
    model="gpt-5.6-sol",
    instructions="You are a helpful assistant.",
    input="Design an algorithm to solve this optimization problem: ...",
)

print(response.output_text)
  • messagesinput
  • system message → instructions or a developer item
  • choices[0].message.contentresponse.output_text
  • for tools and multimodal output, inspect response.output by item type.

DeepSeek Models Reasoning Settings

DeepSeek offers reasoning capabilities through the V4 flagship family (deepseek-v4-pro and deepseek-v4-flash). The stable deepseek-v4-flash ID now uses the official DeepSeek-V4-Flash-0731 release automatically at the same price, with stronger agentic capabilities and low, high, and max reasoning-effort levels. In thinking mode, supported routes can return provider-specific reasoning_content for continuity/observability alongside the final answer. The legacy aliases deepseek-reasoner (→ deepseek-v4-pro) and deepseek-chat (→ deepseek-v4-flash) continue to work but will be retired on July 24, 2026.

Key Features:

  • reasoning_content: Contains the provider-exposed reasoning trace used for tool-call continuity and observability
  • content: Contains the final answer
  • Reasoning trace: Use the provider-exposed trace for debugging, tool-call continuity, and observability
  • Tool Call Integration: Thinking mode works with function calling
  • Thinking Mode Toggle: Use extra_body={"thinking": {"type": "enabled"}} or {"type": "disabled"} where the selected V4 route exposes the toggle
  • Reasoning Effort: DeepSeek-V4-Flash-0731 supports reasoning_effort: "low", "high", or "max"; verify V4-Pro values separately before sharing configuration across the family
  • Context Window: V4 flagships default to a 1M-token context window

Basic Usage (thinking mode, V4-Pro):

python
response = client.chat.completions.create(
    model="deepseek-v4-pro",
    messages=[
        {
            "role": "user",
            "content": "9.11 and 9.8, which is greater?",
        }
    ],
    reasoning_effort="high",
    extra_body={"thinking": {"type": "enabled"}},
)

# Access the reasoning process
reasoning = response.choices[0].message.reasoning_content
# Access the final answer
answer = response.choices[0].message.content
Responses API version This version uses `gpt-5.5` because `deepseek-v4-pro` may not be enabled for `/v1/responses` in the current AvalAI model data.

Use this version when the selected model supports /v1/responses. messages moves to input, and the final text is read from response.output_text.

python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AVALAI_API_KEY"],
    base_url="https://api.avalai.ir/v1",
)

response = client.responses.create(
    model="gpt-5.6-sol",
    instructions="You are a helpful assistant.",
    input="9.11 and 9.8, which is greater?",
)

print(response.output_text)
  • messagesinput
  • system message → instructions or a developer item
  • choices[0].message.contentresponse.output_text
  • for tools and multimodal output, inspect response.output by item type.

Fast, Economical Reasoning (V4-Flash):

python
response = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=[
        {
            "role": "user",
            "content": "Summarize the trade-offs between thinking and non-thinking modes.",
        }
    ],
    extra_body={"thinking": {"type": "enabled"}},
)
Responses API version This version uses `gpt-5.5` because `deepseek-v4-flash` may not be enabled for `/v1/responses` in the current AvalAI model data.

Use this version when the selected model supports /v1/responses. messages moves to input, and the final text is read from response.output_text.

python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AVALAI_API_KEY"],
    base_url="https://api.avalai.ir/v1",
)

response = client.responses.create(
    model="gpt-5.6-sol",
    instructions="You are a helpful assistant.",
    input="Summarize the trade-offs between thinking and non-thinking modes.",
)

print(response.output_text)
  • messagesinput
  • system message → instructions or a developer item
  • choices[0].message.contentresponse.output_text
  • for tools and multimodal output, inspect response.output by item type.

⚠️ Critical: Tool Calls with Thinking Mode

When using tool calls with DeepSeek V4 models in thinking mode (including deepseek-v4-pro, deepseek-v4-flash with thinking enabled, or the legacy deepseek-reasoner alias), you must pass the reasoning_content back to the API in subsequent requests within the same turn. Failure to do so will result in an error:

Missing reasoning_content field in the assistant message

Correct Tool Call Implementation:

python
# When the model returns tool_calls, include reasoning_content in the assistant message
assistant_message = {
    "role": "assistant",
    "content": message.content or "",
    "tool_calls": [...],
    "reasoning_content": message.reasoning_content,  # CRITICAL: Must include this
}
messages.append(assistant_message)

Multi-turn Conversation Rules:

  1. Within a single turn (while processing tool calls): Always include reasoning_content
  2. Between turns (new user message): Only pass content, not reasoning_content

For complete documentation with examples in multiple languages, see the DeepSeek Models Documentation.

Official Reference: DeepSeek Thinking Mode - Tool Calls

Direct HTTP Requests

When making direct HTTP requests or using curl, you can include these parameters directly in the request body:

bash
curl https://api.avalai.ir/v1/chat/completions \
  -H "Authorization: Bearer $AVALAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
"model": "gemini-3.1-flash-lite-preview",
"messages": [{"role": "user", "content": "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ..."}],
"thinking": {"type": "enabled", "budget_tokens": 2000}
}'
Responses API version This version uses `gpt-5.5` because `gemini-3.1-flash-lite-preview` may not be enabled for `/v1/responses` in the current AvalAI model data.

Use this version when the selected model supports /v1/responses. messages moves to input, and the final text is read from response.output_text.

bash
curl https://api.avalai.ir/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AVALAI_API_KEY" \
  -d '
  {
    "model": "gpt-5.6-sol",
    "input": "Solve this complex math problem. Return the final answer, key assumptions, and a concise verification checklist: ...",
    "instructions": "You are a helpful assistant."
  }'
  • messagesinput
  • system message → instructions or a developer item
  • choices[0].message.contentresponse.output_text
  • for tools and multimodal output, inspect response.output by item type.

By adjusting these parameters, you can balance reasoning depth against cost and speed for your specific use case across different model providers.

Managing the Context Window

Ensure sufficient space remains in the context window for both the expected output and the internal reasoning tokens. Complex problems might require thousands or even tens of thousands of reasoning tokens.

You can often find the breakdown of token usage (including reasoning tokens, if exposed by the API) in the usage object of the API response. For OpenAI Responses, output_tokens is the generated-output total including reasoning tokens; the nested reasoning_tokens value is not an additional amount to add.

Example OpenAI Responses usage object:

json
{
  "usage": {
    "input_tokens": 75,
    "input_tokens_details": {
      "cached_tokens": 0
    },
    "output_tokens": 1186,
    "output_tokens_details": {
      "reasoning_tokens": 1024
    },
    "total_tokens": 1261
  }
}

Check the AvalAI Models documentation for context window lengths for specific models.

Use usage telemetry as a tuning loop:

  1. Log model, reasoning.effort, max_output_tokens, latency, status, and token usage for each eval case.
  2. Watch for high reasoning_tokens with low answer-quality gains; lower effort or simplify the prompt when this happens.
  3. Watch for status: "incomplete" or missing visible output; increase max_output_tokens, split the task, or reduce retrieved context.
  4. Keep separate baselines for Chat Completions and Responses, because Responses can preserve reasoning items across tool calls while Chat Completions remains stateless.

Controlling Costs

To manage costs:

  1. Be mindful of the reasoning.effort (or equivalent) parameter if available. Higher effort usually means more tokens.
  2. Use the max_tokens (or equivalent parameter like max_output_tokens) in your API request to limit the total number of tokens generated (reasoning + output).

Allocating Space for Reasoning

max_output_tokens (Responses), max_completion_tokens (Chat Completions), and legacy max_tokens are shared generation budgets on reasoning-capable models: hidden reasoning tokens and visible answer tokens can both consume them. They do not reserve a separate allowance for the final answer.

If reasoning consumes the entire limit, the response can contain a reasoning item but no text. This is a token-budget exhaustion case, not necessarily a content-filter or model-availability problem. Typical evidence includes:

  • Responses: status: "incomplete" and incomplete_details.reason: "max_output_tokens".
  • Usage: output_tokens_details.reasoning_tokens is close to the total output_tokens, while visible text is empty or missing.
  • Chat Completions: finish_reason: "length", possibly before useful visible content is produced.

To recover, increase the applicable output limit within the model's supported maximum, lower reasoning.effort to low or none when the model supports it, simplify or split the task, and leave measured headroom for the final answer. Do not size the limit only from the desired visible answer length. Record usage by model and prompt class because reasoning demand varies between requests, even for the same model.

Handling Incomplete Responses (Example)

Your code should check for indicators that generation was stopped due to token limits.

python
import json

# ... (client setup and initial request as before) ...

try:
    response = client.chat.completions.create(
        model="gpt-5.6-sol",
        messages=[{"role": "user", "content": prompt}],
        max_tokens=300,  # Limit total generated tokens (reasoning + output)
        # Add reasoning parameters if applicable
    )

    finish_reason = response.choices[0].finish_reason
    output_text = response.choices[0].message.content

    if finish_reason == "length":  # Standard finish reason for max_tokens
        print("Ran out of tokens (max_tokens reached).")
        if output_text:
            print("Partial output:", output_text)
        else:
            # This implies the limit was hit during internal reasoning
            print("Ran out of tokens during reasoning phase.")
    elif finish_reason == "stop":
        print("Completed successfully:")
        print(output_text)
    else:
        print(f"Finished with reason: {finish_reason}")
    if output_text:
        print("Output:", output_text)

except Exception as e:
    print(f"An API error occurred: {e}")
javascript
// ... (client setup and initial request as before) ...

async function runReasoningWithLimit() {
  try {
    const response = await client.chat.completions.create({
      model: "gpt-5.6-sol",
      messages: [{ role: "user", content: prompt }],
      max_tokens: 300, // Limit total generated tokens
      // Add reasoning parameters if applicable
    });

    const finish_reason = response.choices[0].finish_reason;
    const output_text = response.choices[0].message.content;

    if (finish_reason === "length") {
      // Standard finish reason for max_tokens
      console.log("Ran out of tokens (max_tokens reached).");
      if (output_text) {
        console.log("Partial output:", output_text);
      } else {
        console.log("Ran out of tokens during reasoning phase.");
      }
    } else if (finish_reason === "stop") {
      console.log("Completed successfully:");
      console.log(output_text);
    } else {
      console.log(`Finished with reason: ${finish_reason}`);
      if (output_text) {
        console.log("Output:", output_text);
      }
    }
  } catch (error) {
    console.error("An API error occurred:", error);
  }
}

runReasoningWithLimit();
go
package main

import (
	"context"
	"fmt"
	"os"

	openai "github.com/openai/openai-go"
)

func main() {
	// ... (client setup as before) ...

	prompt := "..."  // Your prompt here
	maxTokens := 300 // Define max_tokens

	resp, err := client.CreateChatCompletion(
		context.Background(),
		openai.ChatCompletionRequest{
			Model: "gpt-5.5",
			Messages: []openai.ChatCompletionMessage{
				{Role: openai.ChatMessageRoleUser, Content: prompt},
			},
			MaxTokens: maxTokens, // Limit total generated tokens
			// Add reasoning parameters if applicable
		},
	)

	if err != nil {
		fmt.Printf("ChatCompletion error: %v\n", err)
		return
	}

	finishReason := resp.Choices[0].FinishReason
	outputText := resp.Choices[0].Message.Content

	if finishReason == openai.FinishReasonLength { // Check for length finish reason
		fmt.Println("Ran out of tokens (max_tokens reached).")
		if outputText != "" {
			fmt.Println("Partial output:", outputText)
		} else {
			fmt.Println("Ran out of tokens during reasoning phase.")
		}
	} else if finishReason == openai.FinishReasonStop {
		fmt.Println("Completed successfully:")
		fmt.Println(outputText)
	} else {
		fmt.Printf("Finished with reason: %s\n", finishReason)
		if outputText != "" {
			fmt.Println("Output:", outputText)
		}
	}
}
php
<?php
require 'vendor/autoload.php';

// ... (client setup as before) ...

$prompt = "..."; // Your prompt here
$maxTokens = 300; // Define max_tokens

try {
 $response = $client->chat()->create([
 'model' => 'gpt-5.5',
 'messages' => [
 ['role' => 'user', 'content' => $prompt],
 ],
 'max_tokens' => $maxTokens, // Limit total generated tokens
 // Add reasoning parameters if applicable
 ]);

 $finishReason = $response->choices[0]->finishReason;
 // Ensure content exists before accessing
 $outputText = $response->choices[0]->message->content ?? null;

 if ($finishReason === 'length') { // Check for length finish reason
 echo "Ran out of tokens (max_tokens reached).\n";
 if ($outputText) {
 echo "Partial output: " . $outputText . "\n";
 } else {
 echo "Ran out of tokens during reasoning phase.\n";
 }
 } elseif ($finishReason === 'stop') {
 echo "Completed successfully:\n";
 echo $outputText . "\n";
 } else {
 echo "Finished with reason: " . $finishReason . "\n";
 if ($outputText) {
 echo "Output: " . $outputText . "\n";
 }
 }

} catch (Exception $e) {
 echo "An API error occurred: " . $e->getMessage() . "\n";
}
?>
Responses API version

Use this version when the selected model supports /v1/responses. Chat Completions reports token exhaustion with finish_reason: "length"; Responses reports it with status: "incomplete" and incomplete_details.reason: "max_output_tokens".

python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AVALAI_API_KEY"],
    base_url="https://api.avalai.ir/v1",
)

prompt = """
Write a bash script that takes a matrix represented as a string with
format '[1,2],[3,4],[5,6]' and prints the transpose in the same format.
"""

response = client.responses.create(
    model="gpt-5.6-sol",
    reasoning={"effort": "medium"},
    input=[{"role": "user", "content": prompt}],
    max_output_tokens=300,
)

if (
    response.status == "incomplete"
    and response.incomplete_details.reason == "max_output_tokens"
):
    print("Ran out of tokens.")
    if response.output_text:
        print("Partial output:", response.output_text)
    else:
        print("Ran out of tokens during reasoning.")
else:
    print(response.output_text)
javascript
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.AVALAI_API_KEY,
  baseURL: "https://api.avalai.ir/v1",
});

const prompt = `
Write a bash script that takes a matrix represented as a string with
format '[1,2],[3,4],[5,6]' and prints the transpose in the same format.
`;

const response = await client.responses.create({
  model: "gpt-5.6-sol",
  reasoning: { effort: "medium" },
  input: [{ role: "user", content: prompt }],
  max_output_tokens: 300,
});

if (
  response.status === "incomplete" &&
  response.incomplete_details?.reason === "max_output_tokens"
) {
  console.log("Ran out of tokens.");
  if (response.output_text) {
    console.log("Partial output:", response.output_text);
  } else {
    console.log("Ran out of tokens during reasoning.");
  }
} else {
  console.log(response.output_text);
}
bash
PROMPT='Write a bash script that takes a matrix represented as a string with format "[1,2],[3,4],[5,6]" and prints the transpose in the same format.'

jq -n --arg prompt "$PROMPT" '{
  model: "gpt-5.6-sol",
  reasoning: {effort: "medium"},
  input: [{role: "user", content: $prompt}],
  max_output_tokens: 300
}' | curl https://api.avalai.ir/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AVALAI_API_KEY" \
  -d @- | jq '{
    status,
    incomplete_reason: .incomplete_details.reason,
    output_text
  }'
  • messagesinput
  • system message → instructions or a developer item
  • max_tokensmax_output_tokens
  • choices[0].finish_reason == "length"status == "incomplete" with incomplete_details.reason
  • choices[0].message.contentresponse.output_text
  • for tools and multimodal output, inspect response.output by item type.

Advice on Prompting

Prompting reasoning models can differ slightly from prompting standard GPT models.

  • Reasoning Models: Often perform well with higher-level goals and fewer prescribed intermediate steps. Think of them as senior collaborators you can trust to figure out the details.
  • GPT Models: Often benefit from very precise instructions and clear definitions of the desired output format. Think of them as junior collaborators needing explicit guidance.

Experiment with providing the overall objective and letting the reasoning model determine the best path, versus providing detailed steps.

Prompt Examples

(Note: The following examples use current reasoning-capable models such as gpt-5.5, deepseek-v4-pro, qwen3.7-max, and glm-5.2. Adjust parameters/endpoints as needed based on AvalAI's specific implementation.)

1. Coding (Refactoring)

Task: Refactor a React component to change text color based on data.

javascript
// --- Calling Code (Node.js) ---
import OpenAI from "openai";
import * as dotenv from "dotenv";
dotenv.config();

const client = new OpenAI({
  apiKey: process.env.AVALAI_API_KEY,
  baseURL: "https://api.avalai.ir/v1",
});

// Note the change here: Use indentation for the inner code example
// instead of triple backticks within the prompt string.
const prompt = `

Instructions:- Given the React component below, change it so that nonfiction books have red text.
- Return only the refactored React component code in your reply.
- Do not include explanations or markdown code blocks.
- Use four spaces for indentation.
- Keep lines under 80 columns.

Original Code:

    const books = [
     { title: 'Dune', category: 'fiction', id: 1 },
     { title: 'Frankenstein', category: 'fiction', id: 2 },
     { title: 'Moneyball', category: 'nonfiction', id: 3 },
    ];

    export default function BookList() {
     const listItems = books.map(book =>
     <li>
     {book.title}
     </li>
     );

     return (
     <ul>{listItems}</ul>
     );
    }

`.trim();

async function refactorCode() {
  try {
    const response = await client.chat.completions.create({
      model: "gpt-5.6-sol", // Use a suitable reasoning model from AvalAI
      messages: [{ role: "user", content: prompt }],
      temperature: 0.1, // Lower temperature for more predictable code output
    });
    console.log(response.choices[0].message.content);
  } catch (error) {
    console.error("API Error:", error);
  }
}

refactorCode();
python
# --- Calling Code (Python) ---
from openai import OpenAI
import os

client = OpenAI(
    api_key=os.environ.get("AVALAI_API_KEY"),
    base_url="https://api.avalai.ir/v1",
)

# Note the change here: Use indentation for the inner code example
# instead of triple backticks within the prompt string.
prompt = """

Instructions:- Given the React component below, change it so that nonfiction books have red text.
- Return only the refactored React component code in your reply.
- Do not include explanations or markdown code blocks.
- Use four spaces for indentation.
- Keep lines under 80 columns.

Original Code:

    const books = [
     { title: 'Dune', category: 'fiction', id: 1 },
     { title: 'Frankenstein', category: 'fiction', id: 2 },
     { title: 'Moneyball', category: 'nonfiction', id: 3 },
    ];

    export default function BookList() {
     const listItems = books.map(book =>
     <li>
     {book.title}
     </li>
     );

     return (
     <ul>{listItems}</ul>
     );
    }

""".strip()

try:
    response = client.chat.completions.create(
        model="gpt-5.6-sol",  # Use a suitable reasoning model from AvalAI
        messages=[{"role": "user", "content": prompt}],
        temperature=0.1,  # Lower temperature for more predictable code output
    )
    print(response.choices[0].message.content)
except Exception as e:
    print(f"API Error: {e}")
bash
# --- Calling Code (Bash/cURL) ---
# Note the change here: Use indentation for the inner code example
# instead of triple backticks within the prompt string.
PROMPT=$(
  cat <<'EOF'

Instructions:- Given the React component below, change it so that nonfiction books have red text.
- Return only the refactored React component code in your reply.
- Do not include explanations or markdown code blocks.
- Use four spaces for indentation.
- Keep lines under 80 columns.

Original Code:

    const books = [
     { title: 'Dune', category: 'fiction', id: 1 },
     { title: 'Frankenstein', category: 'fiction', id: 2 },
     { title: 'Moneyball', category: 'nonfiction', id: 3 },
    ];

    export default function BookList() {
     const listItems = books.map(book =>
     <li>
     {book.title}
     </li>
     );

     return (
     <ul>{listItems}</ul>
     );
    }

EOF
)

# Escape JSON special characters in the prompt
JSON_PROMPT=$(echo "$PROMPT" | jq -Rsa .)

curl https://api.avalai.ir/v1/chat/completions \
  -H "Authorization: Bearer $AVALAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
 "model": "gpt-5.6-sol",
 "messages": [{"role": "user", "content": '"$JSON_PROMPT"'}],
 "temperature": 0.1
 }'
go
// --- Calling Code (Go) ---
package main

import (
	"context"
	"fmt"
	"os"
	"strings"

	openai "github.com/openai/openai-go"
)

func main() {
	apiKey := os.Getenv("AVALAI_API_KEY")
	baseURL := "https://api.avalai.ir/v1"

	config := openai.DefaultConfig(apiKey)
	config.BaseURL = baseURL
	client := openai.NewClientWithConfig(config)

	// Note the change here: Use indentation for the inner code example
	// instead of triple backticks within the prompt string.
	prompt := strings.TrimSpace(`

Instructions:- Given the React component below, change it so that nonfiction books have red text.
- Return only the refactored React component code in your reply.
- Do not include explanations or markdown code blocks.
- Use four spaces for indentation.
- Keep lines under 80 columns.

Original Code:

    const books = [
     { title: 'Dune', category: 'fiction', id: 1 },
     { title: 'Frankenstein', category: 'fiction', id: 2 },
     { title: 'Moneyball', category: 'nonfiction', id: 3 },
    ];

    export default function BookList() {
     const listItems = books.map(book =>
     <li>
     {book.title}
     </li>
     );

     return (
     <ul>{listItems}</ul>
     );
    }

`)

	temp := float32(0.1)
	resp, err := client.CreateChatCompletion(
		context.Background(),
		openai.ChatCompletionRequest{
			Model: "gpt-5.5", // Use a suitable reasoning model from AvalAI
			Messages: []openai.ChatCompletionMessage{
				{Role: openai.ChatMessageRoleUser, Content: prompt},
			},
			Temperature: &temp,
		},
	)

	if err != nil {
		fmt.Printf("API Error: %v\n", err)
		return
	}
	fmt.Println(resp.Choices[0].Message.Content)
}
php
// --- Calling Code (PHP) ---
<?php
require 'vendor/autoload.php';

use OpenAI\Client;

$apiKey = getenv('AVALAI_API_KEY');
$baseURL = 'https://api.avalai.ir/v1';

// Configure client (example)
$client = OpenAI::client($apiKey);
// Set base URL if needed via factory/config

// Note the change here: Use indentation for the inner code example
// instead of triple backticks within the prompt string.
$prompt = trim(<<<PROMPT

Instructions:- Given the React component below, change it so that nonfiction books have red text.
- Return only the refactored React component code in your reply.
- Do not include explanations or markdown code blocks.
- Use four spaces for indentation.
- Keep lines under 80 columns.

Original Code:

    const books = [
     { title: 'Dune', category: 'fiction', id: 1 },
     { title: 'Frankenstein', category: 'fiction', id: 2 },
     { title: 'Moneyball', category: 'nonfiction', id: 3 },
    ];

    export default function BookList() {
     const listItems = books.map(book =>
     <li>
     {book.title}
     </li>
     );

     return (
     <ul>{listItems}</ul>
     );
    }

PROMPT);


try {
 $response = $client->chat()->create([
 'model' => 'gpt-5.5', // Use a suitable reasoning model from AvalAI
 'messages' => [
 ['role' => 'user', 'content' => $prompt],
 ],
 'temperature' => 0.1,
 ]);

 echo $response->choices[0]->message->content;

} catch (Exception $e) {
 echo "API Error: " . $e->getMessage() . "\n";
}
?>
Responses API version

Use this version when the selected model supports /v1/responses. It keeps the same refactoring task, moves the application rules into instructions, and uses a low reasoning effort because the edit is constrained and well specified.

python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AVALAI_API_KEY"],
    base_url="https://api.avalai.ir/v1",
)

prompt = """
Given the React component below, change it so that nonfiction books have red text.
Return only the refactored React component code in your reply.
Do not include explanations or markdown code blocks.
Use four spaces for indentation.
Keep lines under 80 columns.

Original Code:

    const books = [
     { title: 'Dune', category: 'fiction', id: 1 },
     { title: 'Frankenstein', category: 'fiction', id: 2 },
     { title: 'Moneyball', category: 'nonfiction', id: 3 },
    ];

    export default function BookList() {
     const listItems = books.map(book =>
     <li>
     {book.title}
     </li>
     );

     return (
     <ul>{listItems}</ul>
     );
    }
"""

response = client.responses.create(
    model="gpt-5.6-sol",
    reasoning={"effort": "low"},
    instructions=(
        "Formatting re-enabled\n"
        "You refactor code precisely. Return only the requested code."
    ),
    input=prompt,
)

print(response.output_text)
javascript
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.AVALAI_API_KEY,
  baseURL: "https://api.avalai.ir/v1",
});

const prompt = `
Given the React component below, change it so that nonfiction books have red text.
Return only the refactored React component code in your reply.
Do not include explanations or markdown code blocks.
Use four spaces for indentation.
Keep lines under 80 columns.

Original Code:

    const books = [
     { title: 'Dune', category: 'fiction', id: 1 },
     { title: 'Frankenstein', category: 'fiction', id: 2 },
     { title: 'Moneyball', category: 'nonfiction', id: 3 },
    ];

    export default function BookList() {
     const listItems = books.map(book =>
     <li>
     {book.title}
     </li>
     );

     return (
     <ul>{listItems}</ul>
     );
    }
`;

const response = await client.responses.create({
  model: "gpt-5.6-sol",
  reasoning: { effort: "low" },
  instructions:
    "Formatting re-enabled\nYou refactor code precisely. Return only the requested code.",
  input: prompt,
});

console.log(response.output_text);
bash
PROMPT=$(
  cat <<'EOF'
Given the React component below, change it so that nonfiction books have red text.
Return only the refactored React component code in your reply.
Do not include explanations or markdown code blocks.
Use four spaces for indentation.
Keep lines under 80 columns.

Original Code:

    const books = [
     { title: 'Dune', category: 'fiction', id: 1 },
     { title: 'Frankenstein', category: 'fiction', id: 2 },
     { title: 'Moneyball', category: 'nonfiction', id: 3 },
    ];

    export default function BookList() {
     const listItems = books.map(book =>
     <li>
     {book.title}
     </li>
     );

     return (
     <ul>{listItems}</ul>
     );
    }
EOF
)

jq -n --arg prompt "$PROMPT" '{
  model: "gpt-5.6-sol",
  reasoning: {effort: "low"},
  instructions: "Formatting re-enabled\nYou refactor code precisely. Return only the requested code.",
  input: $prompt
}' | curl https://api.avalai.ir/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AVALAI_API_KEY" \
  -d @-
  • messagesinput
  • system message → instructions or a developer item
  • choices[0].message.contentresponse.output_text
  • for tools and multimodal output, inspect response.output by item type.

2. Coding (Planning)

Task: Plan and generate code for a simple Python Q&A application.

javascript
// --- Calling Code (Node.js) ---
import OpenAI from "openai";
import * as dotenv from "dotenv";
dotenv.config();

const client = new OpenAI({
  apiKey: process.env.AVALAI_API_KEY,
  baseURL: "https://api.avalai.ir/v1",
});

const prompt = `
I want to build a Python app that takes user questions and looks them up
in a simple key-value store (like a dictionary or JSON file) where they
are mapped to answers. If there is a close match (case-insensitive check),
it retrieves the matched answer. If there isn't, it asks the user to
provide an answer and stores the new question/answer pair.

1. Make a plan for the directory structure (e.g., main script, data file).
2. Return the full Python code for the main script.
3. Return an example JSON structure for the data file.
4. Only supply explanatory text at the very beginning and very end, not mixed within the code or file structure output.
`.trim();

async function planProject() {
  try {
    const response = await client.chat.completions.create({
      model: "deepseek-v4-pro", // Use a suitable reasoning model from AvalAI
      messages: [{ role: "user", content: prompt }],
    });
    console.log(response.choices[0].message.content);
  } catch (error) {
    console.error("API Error:", error);
  }
}

planProject();
python
# --- Calling Code (Python) ---
from openai import OpenAI
import os

client = OpenAI(
    api_key=os.environ.get("AVALAI_API_KEY"),
    base_url="https://api.avalai.ir/v1",
)

prompt = """
I want to build a Python app that takes user questions and looks them up
in a simple key-value store (like a dictionary or JSON file) where they
are mapped to answers. If there is a close match (case-insensitive check),
it retrieves the matched answer. If there isn't, it asks the user to
provide an answer and stores the new question/answer pair.

1. Make a plan for the directory structure (e.g., main script, data file).
2. Return the full Python code for the main script.
3. Return an example JSON structure for the data file.
4. Only supply explanatory text at the very beginning and very end, not mixed within the code or file structure output.
""".strip()

try:
    response = client.chat.completions.create(
        model="deepseek-v4-pro",  # Use a suitable reasoning model from AvalAI
        messages=[{"role": "user", "content": prompt}],
    )
    print(response.choices[0].message.content)
except Exception as e:
    print(f"API Error: {e}")
bash
# --- Calling Code (Bash/cURL) ---
PROMPT=$(
  cat <<'EOF'
I want to build a Python app that takes user questions and looks them up
in a simple key-value store (like a dictionary or JSON file) where they
are mapped to answers. If there is a close match (case-insensitive check),
it retrieves the matched answer. If there isn't, it asks the user to
provide an answer and stores the new question/answer pair.

1. Make a plan for the directory structure (e.g., main script, data file).
2. Return the full Python code for the main script.
3. Return an example JSON structure for the data file.
4. Only supply explanatory text at the very beginning and very end, not mixed within the code or file structure output.
EOF
)

# Escape JSON special characters
JSON_PROMPT=$(echo "$PROMPT" | jq -Rsa .)

curl https://api.avalai.ir/v1/chat/completions \
  -H "Authorization: Bearer $AVALAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
 "model": "deepseek-v4-pro",
 "messages": [{"role": "user", "content": '"$JSON_PROMPT"'}]
 }'
go
// --- Calling Code (Go) ---
package main

import (
	"context"
	"fmt"
	"os"
	"strings"

	openai "github.com/openai/openai-go"
)

func main() {
	apiKey := os.Getenv("AVALAI_API_KEY")
	baseURL := "https://api.avalai.ir/v1"

	config := openai.DefaultConfig(apiKey)
	config.BaseURL = baseURL
	client := openai.NewClientWithConfig(config)

	prompt := strings.TrimSpace(`
I want to build a Python app that takes user questions and looks them up
in a simple key-value store (like a dictionary or JSON file) where they
are mapped to answers. If there is a close match (case-insensitive check),
it retrieves the matched answer. If there isn't, it asks the user to
provide an answer and stores the new question/answer pair.

1. Make a plan for the directory structure (e.g., main script, data file).
2. Return the full Python code for the main script.
3. Return an example JSON structure for the data file.
4. Only supply explanatory text at the very beginning and very end, not mixed within the code or file structure output.
`)

	resp, err := client.CreateChatCompletion(
		context.Background(),
		openai.ChatCompletionRequest{
			Model: "deepseek-v4-pro", // Use a suitable reasoning model from AvalAI
			Messages: []openai.ChatCompletionMessage{
				{Role: openai.ChatMessageRoleUser, Content: prompt},
			},
		},
	)

	if err != nil {
		fmt.Printf("API Error: %v\n", err)
		return
	}
	fmt.Println(resp.Choices[0].Message.Content)
}
php
// --- Calling Code (PHP) ---
<?php
require 'vendor/autoload.php';

use OpenAI\Client;

$apiKey = getenv('AVALAI_API_KEY');
$baseURL = 'https://api.avalai.ir/v1';

// Configure client (example)
$client = OpenAI::client($apiKey);
// Set base URL if needed via factory/config

$prompt = trim(<<<PROMPT
I want to build a Python app that takes user questions and looks them up
in a simple key-value store (like a dictionary or JSON file) where they
are mapped to answers. If there is a close match (case-insensitive check),
it retrieves the matched answer. If there isn't, it asks the user to
provide an answer and stores the new question/answer pair.

1. Make a plan for the directory structure (e.g., main script, data file).
2. Return the full Python code for the main script.
3. Return an example JSON structure for the data file.
4. Only supply explanatory text at the very beginning and very end, not mixed within the code or file structure output.
PROMPT);

try {
 $response = $client->chat()->create([
 'model' => 'deepseek-v4-pro', // Use a suitable reasoning model from AvalAI
 'messages' => [
 ['role' => 'user', 'content' => $prompt],
 ],
 ]);

 echo $response->choices[0]->message->content;

} catch (Exception $e) {
 echo "API Error: " . $e->getMessage() . "\n";
}
?>
Responses API version This version uses `gpt-5.5` because `deepseek-v4-pro` may not be enabled for `/v1/responses` in the current AvalAI model data.

Use this version when the selected model supports /v1/responses. It keeps the same planning task and uses medium reasoning effort because the model must design a small structure, write code, and satisfy a strict output contract.

python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AVALAI_API_KEY"],
    base_url="https://api.avalai.ir/v1",
)

prompt = """
I want to build a Python app that takes user questions and looks them up
in a simple key-value store (like a dictionary or JSON file) where they
are mapped to answers. If there is a close match (case-insensitive check),
it retrieves the matched answer. If there isn't, it asks the user to
provide an answer and stores the new question/answer pair.

1. Make a plan for the directory structure (e.g., main script, data file).
2. Return the full Python code for the main script.
3. Return an example JSON structure for the data file.
4. Only supply explanatory text at the very beginning and very end, not mixed within the code or file structure output.
"""

response = client.responses.create(
    model="gpt-5.6-sol",
    reasoning={"effort": "medium"},
    instructions=(
        "Formatting re-enabled\n"
        "You are a senior Python engineer. Follow the requested output contract exactly."
    ),
    input=prompt,
)

print(response.output_text)
javascript
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.AVALAI_API_KEY,
  baseURL: "https://api.avalai.ir/v1",
});

const prompt = `
I want to build a Python app that takes user questions and looks them up
in a simple key-value store (like a dictionary or JSON file) where they
are mapped to answers. If there is a close match (case-insensitive check),
it retrieves the matched answer. If there isn't, it asks the user to
provide an answer and stores the new question/answer pair.

1. Make a plan for the directory structure (e.g., main script, data file).
2. Return the full Python code for the main script.
3. Return an example JSON structure for the data file.
4. Only supply explanatory text at the very beginning and very end, not mixed within the code or file structure output.
`;

const response = await client.responses.create({
  model: "gpt-5.6-sol",
  reasoning: { effort: "medium" },
  instructions:
    "Formatting re-enabled\nYou are a senior Python engineer. Follow the requested output contract exactly.",
  input: prompt,
});

console.log(response.output_text);
bash
PROMPT=$(
  cat <<'EOF'
I want to build a Python app that takes user questions and looks them up
in a simple key-value store (like a dictionary or JSON file) where they
are mapped to answers. If there is a close match (case-insensitive check),
it retrieves the matched answer. If there isn't, it asks the user to
provide an answer and stores the new question/answer pair.

1. Make a plan for the directory structure (e.g., main script, data file).
2. Return the full Python code for the main script.
3. Return an example JSON structure for the data file.
4. Only supply explanatory text at the very beginning and very end, not mixed within the code or file structure output.
EOF
)

jq -n --arg prompt "$PROMPT" '{
  model: "gpt-5.6-sol",
  reasoning: {effort: "medium"},
  instructions: "Formatting re-enabled\nYou are a senior Python engineer. Follow the requested output contract exactly.",
  input: $prompt
}' | curl https://api.avalai.ir/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AVALAI_API_KEY" \
  -d @-
  • messagesinput
  • system message → instructions or a developer item
  • choices[0].message.contentresponse.output_text
  • for tools and multimodal output, inspect response.output by item type.

3. STEM Research

Task: Ask for potential compounds for antibiotic research.

javascript
// --- Calling Code (Node.js) ---
import OpenAI from "openai";
import * as dotenv from "dotenv";
dotenv.config();

const client = new OpenAI({
  apiKey: process.env.AVALAI_API_KEY,
  baseURL: "https://api.avalai.ir/v1",
});

const prompt = `
What are three compounds or classes of compounds we should consider
investigating further to advance research into new antibiotics,
especially against resistant bacteria? Briefly explain why each
is promising.
`.trim();

async function researchQuery() {
  try {
    const response = await client.chat.completions.create({
      model: "gemini-3.1-pro-preview", // Use a suitable reasoning model from AvalAI
      messages: [{ role: "user", content: prompt }],
    });
    console.log(response.choices[0].message.content);
  } catch (error) {
    console.error("API Error:", error);
  }
}

researchQuery();
python
# --- Calling Code (Python) ---
from openai import OpenAI
import os

client = OpenAI(
    api_key=os.environ.get("AVALAI_API_KEY"),
    base_url="https://api.avalai.ir/v1",
)

prompt = """
What are three compounds or classes of compounds we should consider
investigating further to advance research into new antibiotics,
especially against resistant bacteria? Briefly explain why each
is promising.
""".strip()

try:
    response = client.chat.completions.create(
        model="gemini-3.1-pro-preview",  # Use a suitable reasoning model from AvalAI
        messages=[{"role": "user", "content": prompt}],
    )
    print(response.choices[0].message.content)
except Exception as e:
    print(f"API Error: {e}")
bash
# --- Calling Code (Bash/cURL) ---
PROMPT=$(
  cat <<'EOF'
What are three compounds or classes of compounds we should consider
investigating further to advance research into new antibiotics,
especially against resistant bacteria? Briefly explain why each
is promising.
EOF
)

# Escape JSON special characters
JSON_PROMPT=$(echo "$PROMPT" | jq -Rsa .)

curl https://api.avalai.ir/v1/chat/completions \
  -H "Authorization: Bearer $AVALAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
 "model": "gemini-3.1-pro-preview",
 "messages": [{"role": "user", "content": '"$JSON_PROMPT"'}]
 }'
go
// --- Calling Code (Go) ---
package main

import (
	"context"
	"fmt"
	"os"
	"strings"

	openai "github.com/openai/openai-go"
)

func main() {
	apiKey := os.Getenv("AVALAI_API_KEY")
	baseURL := "https://api.avalai.ir/v1"

	config := openai.DefaultConfig(apiKey)
	config.BaseURL = baseURL
	client := openai.NewClientWithConfig(config)

	prompt := strings.TrimSpace(`
What are three compounds or classes of compounds we should consider
investigating further to advance research into new antibiotics,
especially against resistant bacteria? Briefly explain why each
is promising.
`)

	resp, err := client.CreateChatCompletion(
		context.Background(),
		openai.ChatCompletionRequest{
			Model: "gemini-3.1-pro-preview", // Use a suitable reasoning model from AvalAI
			Messages: []openai.ChatCompletionMessage{
				{Role: openai.ChatMessageRoleUser, Content: prompt},
			},
		},
	)

	if err != nil {
		fmt.Printf("API Error: %v\n", err)
		return
	}
	fmt.Println(resp.Choices[0].Message.Content)
}
php
// --- Calling Code (PHP) ---
<?php
require 'vendor/autoload.php';

use OpenAI\Client;

$apiKey = getenv('AVALAI_API_KEY');
$baseURL = 'https://api.avalai.ir/v1';

// Configure client (example)
$client = OpenAI::client($apiKey);
// Set base URL if needed via factory/config

$prompt = trim(<<<PROMPT
What are three compounds or classes of compounds we should consider
investigating further to advance research into new antibiotics,
especially against resistant bacteria? Briefly explain why each
is promising.
PROMPT);

try {
 $response = $client->chat()->create([
 'model' => 'gemini-3.1-pro-preview', // Use a suitable reasoning model from AvalAI
 'messages' => [
 ['role' => 'user', 'content' => $prompt],
 ],
 ]);

 echo $response->choices[0]->message->content;

} catch (Exception $e) {
 echo "API Error: " . $e->getMessage() . "\n";
}
?>
Responses API version This version uses `gpt-5.5` because `gemini-3.1-pro-preview` may not be enabled for `/v1/responses` in the current AvalAI model data.

Use this version when the selected model supports /v1/responses. It keeps the same research-style prompt, but adds an explicit safety boundary so the answer stays high level and does not provide synthesis or wet-lab instructions.

python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AVALAI_API_KEY"],
    base_url="https://api.avalai.ir/v1",
)

prompt = """
What are three compounds or classes of compounds we should consider
investigating further to advance research into new antibiotics,
especially against resistant bacteria? Briefly explain why each is promising.
"""

response = client.responses.create(
    model="gpt-5.6-sol",
    reasoning={"effort": "high"},
    instructions=(
        "Formatting re-enabled\n"
        "Answer at a high scientific level. Do not include synthesis steps, "
        "dosages, protocols, or operational wet-lab instructions."
    ),
    input=prompt,
)

print(response.output_text)
javascript
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.AVALAI_API_KEY,
  baseURL: "https://api.avalai.ir/v1",
});

const prompt = `
What are three compounds or classes of compounds we should consider
investigating further to advance research into new antibiotics,
especially against resistant bacteria? Briefly explain why each is promising.
`;

const response = await client.responses.create({
  model: "gpt-5.6-sol",
  reasoning: { effort: "high" },
  instructions:
    "Formatting re-enabled\nAnswer at a high scientific level. Do not include synthesis steps, dosages, protocols, or operational wet-lab instructions.",
  input: prompt,
});

console.log(response.output_text);
bash
PROMPT=$(
  cat <<'EOF'
What are three compounds or classes of compounds we should consider
investigating further to advance research into new antibiotics,
especially against resistant bacteria? Briefly explain why each is promising.
EOF
)

jq -n --arg prompt "$PROMPT" '{
  model: "gpt-5.6-sol",
  reasoning: {effort: "high"},
  instructions: "Formatting re-enabled\nAnswer at a high scientific level. Do not include synthesis steps, dosages, protocols, or operational wet-lab instructions.",
  input: $prompt
}' | curl https://api.avalai.ir/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AVALAI_API_KEY" \
  -d @-
  • messagesinput
  • system message → instructions or a developer item
  • choices[0].message.contentresponse.output_text
  • for tools and multimodal output, inspect response.output by item type.

Use Case Examples

Explore the AvalAI Cookbook (if available) or community resources for more examples of applying reasoning models to tasks like data validation, routine generation, and complex analysis.