Skip to content

Document ingestion for developers

Your documents.
Ready for
your application.

Turn files into searchable content and cited answers. One integration for ingestion, extraction, and retrieval.

REST API TypeScript SDK Streaming chat

Start with a filePOST /v1/ingest/bulk
Upload a document
cURL
curl "https://docslurp.io/v1/ingest/bulk" \
  -H "Authorization: Bearer $DOCSLURP_API_KEY" \
  -F "[email protected]"

Returns a workspace, a run, and your documents. Processing continues asynchronously.

  1. 01UploadYour files
  2. 02ExtractText + structure
  3. 03RetrieveCited results

Add your API key to get started. Setup guide

Built for more than PDFs.

  • PDF
  • DOCX
  • XLSX
  • HTML
  • EPUB
  • Images
  • Audio
  • Video

A small API. A complete workflow.

From a file
to a useful answer.

Keep the ingestion plumbing out of your application. Upload, wait for processing, then query the workspace.

  1. 01

    Send your documents

    Let the upload create a workspace, or pass a workspace ID to add to an existing corpus.

  2. 02

    Follow the run

    Processing happens in the background. The SDK polls until completion and surfaces failures.

  3. 03

    Retrieve with context

    Get relevant passages with citations. Use them in search, RAG, or your own agent workflow.

Explore the SDK
quickstart.ts
TypeScript · Node.js
import { createDocSlurpClient } from '@docslurp/sdk';
import { createReadStream } from 'node:fs';

const client = createDocSlurpClient();

// Upload a file.
const { workspace, run } = await client.upload(createReadStream('handbook.pdf'));

// Processing is asynchronous. Wait before searching.
await client.waitForRun(run.id);

const { results } = await client.search({
  workspaceId: workspace.id,
  q: 'What is the parental leave policy?',
});

for (const hit of results) {
  console.log(hit.text, hit.citation);
}

Server-side example · Set DOCSLURP_API_KEY once; the client reads it automatically.

Built to be inspected

The details matter.
Especially in production.

Know where an answer came from, control how documents are processed, and see what each run consumed.

01Document + page references

Keep the source in reach.

Search results carry document and page citations. Open the source to check the passage behind an answer.

Read the docs
02Optional pipeline stages

Configure each workspace.

Choose extraction, indexing, and enrichment settings for each corpus. Add structured output when your application needs a schema.

Read the docs
03Your application, your UX

Bring your own interface.

Use the REST API, TypeScript client, or OpenAI-compatible chat endpoint. Stream answers into your UI, with a non-streaming option.

Read the docs
04Run-level visibility

See what processing takes.

Inspect run progress, stage timing, token usage, and recorded costs. Follow a file from upload through extraction and indexing.

Read the docs

Build on your own documents

Make your next upload useful.

Try a file in the dashboard, or go straight to the API.