Duct for developers

Search for your app, documents included.

One API to index PDFs, Office files, scans and JSON, and search them by keyword or meaning, on your own servers.

Apache 2.0 · Self-hosted · One Docker image

# 1. A key and a collection
duct keys create --name "my app" --scopes admin
curl -X POST $DUCT/v1/collections -H "Authorization: Bearer $KEY" \
  -d '{"name":"contracts"}'

# 2. Files, or your own text under your own ids
curl -X POST $DUCT/v1/collections/contracts/files \
  -H "Authorization: Bearer $KEY" -F file=@msa.pdf -F id=msa-2026

# 3. Search with filters and facets
curl "$DUCT/v1/collections/contracts/search?q=indemnity&facets=client" \
  -H "Authorization: Bearer $KEY"

Two hard projects, one API

Adding good document search to an app is two hard projects. First, reading PDFs, Office files and scans into clean text with page numbers. Second, running and tuning a search engine, or paying for a hosted one with your data leaving the region.

  • A versioned /v1 REST API with API keys and scopes, and collections: one index per app or per customer.
  • Index files (99 types, scans included) or your own text and JSON under your own ids. Update and delete by id.
  • Search with page numbers, highlighted snippets, metadata filters, facet counts, sorting and pagination.
  • An OpenAPI spec, TypeScript and Python clients, and one Docker image. Apache 2.0.

Built for multi-tenant apps

What you’d otherwise build around a search engine.

Collections: one index per tenant

Each collection is a separate index on disk, so one customer’s documents never appear in another’s results. Keys can be limited to some collections.

Filters and facets

Filter on any metadata field and get value counts for faceted navigation, sorted by relevance or by a field such as year.

Keys with scopes

Create keys with search, write or admin scope from the command line or the API. Only a hash is stored, and each key is rate limited.

A search response

Every hit knows its page, so your users can check it.

{
  "hits": [{
    "id": "msa-2026",
    "title": "Master services agreement",
    "score": 7.42,
    "format": "pdf",
    "page": 12,
    "page_label": "p. 12",
    "snippet": "The indemnity survives termination of this agreement…",
    "highlight": "The <mark>indemnity</mark> survives termination of this agreement…",
    "metadata": { "client": "acme", "year": 2026 }
  }],
  "facets": { "client": { "acme": 14, "globex": 3 } },
  "offset": 0, "limit": 10, "has_more": false, "took_ms": 4
}

Clients

// TypeScript: npm install @docfide/duct
import { DuctClient } from '@docfide/duct/client'

const duct = new DuctClient({ url, key })
await duct.upsert('help', 'refunds', {
  title: 'Refunds', text: '…', metadata: { lang: 'en' },
})
const { hits } = await duct.search('help', {
  q: 'refund', filter: { lang: 'en' },
})
# Python: standard library only
from duct_client import DuctClient

duct = DuctClient(url, key=key)
duct.upload_file("contracts", "msa.pdf", id="msa")
hits = duct.search("contracts", "indemnity",
                   filter={"client": "acme"})["hits"]

Full reference: the developer guide and the OpenAPI spec at /v1/openapi.json on your server.

Run it on your own servers today.