Chat completions (OpenAI-compatible)
POST /v1/chat/completionsOpenAI-compatible chat completions, base tier: point any OpenAI SDK at base_url https://agent402.tools/v1 and pay per call over x402 or MPP (the 402 lists every accepted rail), no API key, no signup. Send POST /v1/chat/completions with the required field messages and pay quoted per request from $0.02 over x402 or MPP (there is no free tier). It returns a JSON object with id, object, created, model, choices and 1 more.
Returns the standard chat.completion object (choices[].message.content, usage). Budget/mid models: gpt-4o-mini (the default when model is omitted), claude haiku, gemini flash, deepseek, llama, mistral, qwen. Full wire compatibility incl. tools/function-calling and response_format. GET /v1/models lists every model. Streaming supported (stream: true). One flat price per call up to the tier caps; /v1/metered/chat/completions quotes each call from its own size instead. Model-backed.
Parameters
| Name | Type | Required | Description |
|---|---|---|---|
model | string | no | Model id - OpenRouter form (openai/gpt-4o-mini) or bare OpenAI form (gpt-4o-mini). GET /v1/models lists the allowlist per tier. Optional: omit it and the tier serves its documented default (x402.defaultModel on /v1/models), named back in agent402_default_model; the price does not change |
messages | array | yes | OpenAI chat messages: [{role, content}] - text and image_url content blocks supported |
max_tokens | number | no | Output token cap (clamped to the tier maximum) |
zdr | boolean | no | Optional - true routes only to zero-data-retention providers (OpenRouter provider.zdr); the only provider preference a caller may set. |
cache_control | any | no | Optional - prompt caching preference. Default ON ({type:"ephemeral"}, 5-minute TTL): repeated prefixes across your turns are served from the provider cache (same price to you). Send false to disable. ttl:"1h" is not offered. |
reasoning | object | no | Optional - {effort: "none"|"minimal"|"low"|"medium"|"high"|"xhigh"|"max", max_tokens?, exclude?, enabled?}. Reasoning tokens count against max_tokens. Omitted: low effort on the budget tiers, the model default on premium. reasoning_effort (string) is accepted as an alias. |
max_completion_tokens | integer | no | Optional - alias of max_tokens (newer OpenAI SDKs send this). |
tools | array | no | Optional - OpenAI function tools {type:"function", function:{...}}, or a tool namespace {type:"namespace", name, tools:[...]} (flattened into its functions). The pro and premium routes also accept the bounded server tools openrouter:web_search, openrouter:web_fetch and openrouter:datetime with server-owned limits (GET /v1/models lists them); stop_server_tools_when and max_tool_calls are refused. A request with a server tool is never served from the prompt cache. |
Example request
curl -i -X POST https://agent402.tools/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"openai/gpt-4o-mini","messages":[{"role":"user","content":"Reply with exactly: OK"}],"max_tokens":5}'
Without payment this returns HTTP 402 Payment Required with the exact price for v1-chat; any x402 v2 or MPP client pays it and retries.
Example response
{
"id": "gen-…",
"object": "chat.completion",
"created": 1750000000,
"model": "openai/gpt-4o-mini",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "OK"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 12,
"completion_tokens": 1,
"total_tokens": 13
}
}
| Field | Type | Always present | In the example |
|---|---|---|---|
id | string | yes | gen-… |
object | string | yes | chat.completion |
created | number | yes | 1750000000 |
model | string | yes | openai/gpt-4o-mini |
choices | array of objects | yes | 1 item in the example |
usage | object | yes | 3 fields: prompt_tokens, completion_tokens, total_tokens |
From an MCP client
catalog.call {
"slug": "v1-chat",
"params": {
"model": "openai/gpt-4o-mini",
"messages": [
{
"role": "user",
"content": "Reply with exactly: OK"
}
],
"max_tokens": 5
}
}
The hosted connector at https://agent402.tools/mcp needs a payment for v1-chat; the stdio package pays it from a wallet or from AGENT402_CREDITS_KEY. Local install: npx -y agent402-mcp.
Errors and behavior
messagesis required. An input the tool rejects returns an HTTP 4xx whose body carrieserror,tool,expected,requiredandexample, so the caller can correct it.- A paid call that ends in any status of 400 or above is not charged: settlement is cancelled when the tool fails.
- Wallet-only: this tool runs a model, so it has no proof-of-work tier. A prepaid card-credits key (
Authorization: Bearer a402_...) also pays it. - Model-backed: the answer is generated by a model, so the same input can produce different wording.
- Priced per request: the 402 quotes this body, between $0.02 and $0.5.
- Flat per call for the models this tier serves. A body naming another flat tier's model (nano, base, pro, premium) is quoted at that tier's price in the 402 and served under that tier; model "auto" is quoted and served as the auto tier. The answer names the tier in agent402_tier. The live 402 is always the price.
- A
GETorHEADto /v1/chat/completions returns the same 402 quote, so the price can be read without a body. - An
Idempotency-Keyheader makes a retried paid call replay the first 200 instead of charging again (streamed responses are not replayed).
Paid call (JavaScript agent)
import { wrapFetchWithPayment } from "@x402/fetch";
import { x402Client } from "@x402/core/client";
import { registerExactEvmScheme } from "@x402/evm/exact/client";
import { privateKeyToAccount } from "viem/accounts";
const client = new x402Client();
client.setSpendControls?.(false); // keep your own spending ceiling in code
registerExactEvmScheme(client, { signer: privateKeyToAccount(KEY) });
const payFetch = wrapFetchWithPayment(fetch, client);
const res = await payFetch("https://agent402.tools/v1/chat/completions", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
"model": "openai/gpt-4o-mini",
"messages": [
{
"role": "user",
"content": "Reply with exactly: OK"
}
],
"max_tokens": 5
}),
});
Related tools
Text-to-speech (OpenAI-compatible)
POST /v1/audio/speechOpenAI-compatible text-to-speech over x402 - point any OpenAI SDK's audio.speech.create() at base_url https://agent402.t…
Chat completions - auto tier (eval-ranked routing)
POST /v1/auto/chat/completionsOpenAI-compatible chat completions with server-side model choice: omit "model" (or send "auto") and the gateway routes t…
Grounded chat (web search, OpenAI-compatible)
POST /v1/grounded/chat/completionsOpenAI-compatible chat completions GROUNDED in a live web search on every call: the gateway runs an Exa search (up to 5 …
Chat completions - metered (pay what the call costs)
POST /v1/metered/chat/completionsOpenAI-compatible chat completions billed per request from what the call costs: the 402 quotes exact-BPE input plus your…
Chat completions - nano tier
POST /v1/nano/chat/completionsOpenAI-compatible chat completions, nano tier: gpt-6-luna (the default), gpt-5.6-luna, gpt-5-nano, gemini flash-lite, sm…
Chat completions - premium tier
POST /v1/premium/chat/completionsOpenAI-compatible chat completions, premium tier: gpt-5, gpt-6 astra, o3 and o4-mini, claude opus, claude fable 5.1 - pa…