Rerank (Cohere-compatible)
POST /v1/rerankRerank documents against a query over x402 - the Cohere /rerank wire ({query, documents[], top_n} -> results with relevance_score), served by cohere/rerank-v3.5, $0.002 per call in USDC, no API key, no signup. Send POST /v1/rerank with the required fields query and documents and pay $0.002 per call over x402 or MPP (there is no free tier). It returns a JSON object with id, model, results and usage.
Up to 50 documents (1,600 chars each, 40k total) and a 500-char query per call. Deterministic, so a byte-identical repeat within 10 minutes is served FREE from cache (X-Cache: hit; opt out with cache:false). The retrieval companion to /v1/embeddings - embed and recall, then rerank the top candidates.
Parameters
| Name | Type | Required | Description |
|---|---|---|---|
query | string | yes | The search query (max 500 chars) Also accepted as q, search, term, keyword, question. |
documents | array of string | yes | Documents to rank (1-50 strings, 1,600 chars each, 40k total) |
top_n | integer | no | Optional - return only the top N results |
cache | boolean | no | Optional - false disables the default-on response cache |
Example request
curl -i -X POST https://agent402.tools/v1/rerank \
-H "Content-Type: application/json" \
-d '{"query":"What is the capital of France?","documents":["Paris is the capital of France.","Berlin is the capital of Germany.","Madrid is in Spain."],"top_n":2}'
Without payment this returns HTTP 402 Payment Required with the exact price for v1-rerank; any x402 v2 or MPP client pays it and retries.
Example response
{
"id": "gen-rerank-…",
"model": "rerank-v3.5",
"results": [
{
"index": 0,
"relevance_score": 0.89,
"document": {
"text": "Paris is the capital of France."
}
},
{
"index": 1,
"relevance_score": 0.15,
"document": {
"text": "Berlin is the capital of Germany."
}
}
],
"usage": {
"search_units": 1
}
}
| Field | Type | Always present | In the example |
|---|---|---|---|
id | string | yes | gen-rerank-… |
model | string | yes | rerank-v3.5 |
results | array of objects | yes | 2 items in the example |
usage | object | yes | 1 field: search_units |
From an MCP client
catalog.call {
"slug": "v1-rerank",
"params": {
"query": "What is the capital of France?",
"documents": [
"Paris is the capital of France.",
"Berlin is the capital of Germany.",
"Madrid is in Spain."
],
"top_n": 2
}
}
The hosted connector at https://agent402.tools/mcp needs a payment for v1-rerank; the stdio package pays it from a wallet or from AGENT402_CREDITS_KEY. Local install: npx -y agent402-mcp.
Errors and behavior
queryanddocumentsare required. An input the tool rejects returns an HTTP 4xx whose body carrieserror,tool,expected,requiredandexample, so the caller can correct it.- A paid call that ends in any status of 400 or above is not charged over x402, MPP or a prepaid credits key: settlement is cancelled when the tool fails. The exception is a Tempo push credential, a transfer the buyer sent before the call: it settles before the tool runs, so if the tool then fails the payment is recorded as a refund owed to the paying wallet.
- Wallet-only: this tool runs a model, so it has no proof-of-work tier. A prepaid card-credits key (
Authorization: Bearer a402_...) also pays it. - Model-backed: the answer is generated by a model, so the same input can produce different wording.
- A
GETorHEADto /v1/rerank returns the same 402 quote, so the price can be read without a body. - An
Idempotency-Keyheader makes a retried paid call replay the first 200 instead of charging again (an answer larger than 1 MB and a streamed response are not replayed).
Paid call (JavaScript agent)
import { wrapFetchWithPayment } from "@x402/fetch";
import { x402Client } from "@x402/core/client";
import { registerExactEvmScheme } from "@x402/evm/exact/client";
import { privateKeyToAccount } from "viem/accounts";
const client = new x402Client();
client.setSpendControls?.(false); // keep your own spending ceiling in code
registerExactEvmScheme(client, { signer: privateKeyToAccount(KEY) });
const payFetch = wrapFetchWithPayment(fetch, client);
const res = await payFetch("https://agent402.tools/v1/rerank", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
"query": "What is the capital of France?",
"documents": [
"Paris is the capital of France.",
"Berlin is the capital of Germany.",
"Madrid is in Spain."
],
"top_n": 2
}),
});
Related tools
Embeddings (OpenAI-compatible)
POST /v1/embeddingsOpenAI-compatible text embeddings over x402 - point any OpenAI SDK at base_url https://agent402.tools/v1 and pay $0.002 …
Text-to-speech (OpenAI-compatible)
POST /v1/audio/speechOpenAI-compatible text-to-speech over x402 - point any OpenAI SDK's audio.speech.create() at base_url https://agent402.t…
Chat completions (OpenAI-compatible)
POST /v1/chat/completionsOpenAI-compatible chat completions, base tier: point any OpenAI SDK at base_url https://agent402.tools/v1 and pay per ca…
Chat completions - auto tier (eval-ranked routing)
POST /v1/auto/chat/completionsOpenAI-compatible chat completions with server-side model choice: omit "model" (or send "auto") and the gateway routes t…
Grounded chat (web search, OpenAI-compatible)
POST /v1/grounded/chat/completionsOpenAI-compatible chat completions GROUNDED in a live web search on every call: the gateway runs an Exa search (up to 5 …
Chat completions - metered (pay what the call costs)
POST /v1/metered/chat/completionsOpenAI-compatible chat completions billed per request from what the call costs: the 402 quotes exact-BPE input plus your…