Duct's goal is that search feels instant, which in practice means results in well under 100 ms. On a laptop-sized library it already was. Then we measured a big one.
Measuring first
We generated a synthetic library of 30,000 documents (90,000 passages) with a realistic spread of common and rare words, and timed each kind of query 25 times, through the library and through the HTTP API the app uses.
| Query | Before | After |
|---|---|---|
| A common word | 114 ms | 28 ms |
| Two words | 119 ms | 36 ms |
| A long query | 154 ms | 61 ms |
| A word that isn't there | 47 ms | 6 ms |
Those are medians on one MacBook. Real libraries are messier than synthetic ones, so read them as "how it scales", not as a promise.
The last row was the clue. A word that appears nowhere should cost almost nothing, yet it took 47 ms. Something was doing work proportional to the size of the library on every search.
Two queries doing work nobody saw
Duct stores its index in SQLite and uses FTS5 for full-text search. A search runs two queries: one over the text, and one over file names, because people often search for a document by what it's called.
The file-name query joined every document to its first passage before checking the name:
SELECT ... FROM documents d
JOIN chunks c ON c.document_id = d.id
AND c.idx = (SELECT min(idx) FROM chunks WHERE document_id = d.id)
WHERE lower(d.display_name) LIKE ?
SQLite's plan scanned all 90,000 passages and ran the subquery for each. Matching the names first, then fetching the first passage of the few that match, takes 5.5 ms instead of 46.
The text query ranked matches with bm25(), but it also asked for each match's excerpt (snippet()), its full text and its document, all in the same statement as ORDER BY ... LIMIT 30. SQLite carries every selected column through the sort, so for a common word it built tens of thousands of excerpts and kept thirty.
Ranking in the full-text index alone took 21 ms. So now Duct does it in two steps:
- Rank in FTS5 and take the best candidates' ids and scores.
- Fetch text, document and excerpt for just those.
Keeping filtered results identical
The catch is filters. You can search one folder, one file type, a date range, or (on a team server) only what you're allowed to see. The old query applied filters before ranking. Ranking first and filtering after could return too few results if the best matches were all outside the filter.
So the filter is checked on the best candidates, and if it leaves fewer than we need, Duct looks further down the ranking: first 4 times as many candidates as results asked for, then 32 times, up to a few thousand. If even that isn't enough, it falls back to the original query that filters while ranking. Results are the same either way; there's a test with 60 strong matches outside a folder and 3 weak ones inside it, which must still find all 3.
"Did you mean", faster and better
When a search finds nothing, Duct suggests a spelling, using only words that are in your documents. It did that by reading every indexed word starting with the same letter, each time. That's now cached until the index changes.
While testing we found it missed everyday misspellings like "tarriff" and "shedule". The index stores word stems ("schedul"), and the comparison cut the typed word to the stem's length, which turned one extra letter into two edits. It now compares the stem against a few lengths of the typed word and keeps the closest.
What's next
Long queries are the slowest case now, at about 60 ms, because every extra word adds matches to rank. That's comfortably under our target, but it's the next thing we'll look at. The benchmark script is in the repository if you want to try it on your own library: npm run bench -- <folder>.