Search and chat API
POST /v1/search retrieves portable passages with structured citations and corpus availability. It returns {results, sessionId, availability} and generates no answer. q and workspaceId are required; limit is an integer from 1 to 100. Queries are limited to 16,000 characters. Filters include documentId, runId, filenameContains, and comma-separated tags or mimeTypes. HyDE, reranking, compression, query rewriting/decomposition, graph expansion, and other optional retrieval enhancements are disabled unless requested. Query embeddings may still incur provider usage; actual usage is logged and persisted.
Batch retrieval
POST /v1/search/batch accepts the same search input objects:
{
"queries": [
{"q": "support hours", "workspaceId": "workspace-id", "limit": 3},
{"q": "urgent incidents", "workspaceId": "workspace-id", "tags": "operations"}
],
"concurrency": 4
}
The batch contains 1–20 queries; concurrency is an integer from 1 to 4, defaulting to 4. The entire envelope is validated before any retrieval starts. The response is {results: [...]}, preserving query order even when requests complete out of order. Each result is either {ok: true, data: {results, sessionId, availability}} or {ok: false, error: {code, message, status}}. A valid batch returns HTTP 200 even when individual items fail. Invalid envelopes return HTTP 400.
Every item independently checks the API key’s search scope and organization/workspace access. Item error codes include FORBIDDEN (403), WORKSPACE_NOT_FOUND (404), INVALID_SEARCH_SCOPE (400), and SEARCH_FAILED (500). Unexpected internal errors are sanitized. Disconnecting prevents queued retrieval from starting; in-flight retrieval may finish and record usage. Unstarted items carry REQUEST_ABORTED (499) if a response remains deliverable.
Grounded chat
POST /v1/chat is the canonical workspace chat endpoint. POST /v1/search/chat remains a compatibility alias with the same validated request and response. Send {q, workspaceId, sessionId?, model?}. sessionId and model are nonempty strings of at most 256 characters; all fields are checked before model calls. The response includes an answer, citations, and the chat session ID. Native clients can keep using GET /v1/search/chat/stream; OpenAI-compatible clients retain /v1/chat/completions and its JSON/SSE fallback.
Independent conversations can run concurrently. Await each reply before reusing its sessionId for the next turn. Simultaneous turns in the same conversation have no ordering guarantee. Session ownership must be verified before a caller-provided session can be reused; a database failure fails closed.
All routes use the existing authenticated request actor. API keys remain scoped to their organization and require search permission. Batch execution uses direct service calls, without internal HTTP calls or forwarding credentials.
Retrieval usage and cost
Successful GET/POST /v1/search responses include usage. Batch responses expose it in each successful item’s data.usage; MCP workspace.search and workspace.search-many return the same fields.
| Field | Meaning |
|---|---|
inputTokens, outputTokens, totalTokens |
Recorded tokens for the retrieval request, including enabled AI enhancements |
costCents |
Recorded AI cost in US cents; preserves fractional cents |
currency |
USD |
costSource |
provider, estimate, mixed, unknown, or none when no AI usage was recorded |
scope |
retrieval |
This is request metadata, not an invoice. It excludes ingestion, storage, and calls to your own model. Provider-reported costs and estimates remain distinguishable; zero recorded usage does not mean every service operation is free. Native chat’s usage describes answer generation.
The TypeScript SDK supports direct for await (const hit of client.search(input)). Save the search and await it after iteration to read usage and sessionId without sending a second request.