Better Vector Search for Long Documents: Chunking Inside Manticore Search
Pangram verdict · v3.3
We believe this text is mainly AI, with some human-written content.
AI likelihood · overall
AIArticle text · 1,584 words · 1 segments analyzed
Say you are building search over your team's internal documentation — guides, runbooks, postmortems. You have a table with auto embeddings : you insert text, Manticore runs the model and fills the vector column for you. (If that is new to you, start with vector search in Manticore .) You load a 4,000-word document. The insert succeeds. The search works. Everything looks fine.Except the model you picked has a 512-token input window, and that document is about 5,000 tokens long. The model read the first 380 words and threw away the other 3,600. Nothing in the document past that point can ever be retrieved, and nothing anywhere told you. The embedding may not represent the document as a whole either.Until now, you would usually split the document into several pieces yourself, create embeddings for each one, and then work out how to combine the results if you wanted document search rather than chunk search. Manticore now handles this in the table definition: add chunk_strategy to the vector column in CREATE TABLE, and Manticore splits each document into chunks, embeds every chunk, and searches all of them:DROP TABLE IF EXISTS docs; CREATE TABLE docs ( title text, content text, chunks float_vector_array knn_type='hnsw' hnsw_similarity='cosine' model_name='Xenova/all-MiniLM-L6-v2' from='title,content' chunk_strategy='sentence' max_tokens='256' overlap_tokens='32' ); That is the whole feature. No ingest pipeline, no splitter library, no second table for chunks, no GROUP BY to fold chunk hits back into documents.TL;DRFive strategies: truncate (the old default), mean, fixed, recursive, sentence. Set with chunk_strategy on a model-backed vector column.truncate and mean produce one vector per document and work on a float_vector column. fixed, recursive and sentence produce many, so they need a float_vector_array column.A document is still one search result. Chunks compete individually, and Manticore returns the document once, with knn_dist() reporting the distance to its closest chunk. k counts documents, not chunks.Tuning knobs: max_tokens (chunk size), overlap_tokens (shared tokens between neighbors), max_chunks (ceiling per document).Measured on the Manticore manual (189 pages, ~298k words): for content buried past the model's window, recall@5 went from 55.1% → 83.3% and MRR from 0.44 → 0.70, at ~2.5× the RAM and ~4× the ingest time.Queries are never chunked. A query is short enough to embed as a whole; only stored documents are split.The problem, shown with a small exampleSuppose you have four documents:Backup and restore runbook — about 700 words, roughly 900 tokens. Backup schedules, retention, restore drills, credentials, capacity planning. The last section explains how to rotate the TLS certificate used by the replication port.Monitoring and alerting guide — unrelated.Getting started with the CLI — unrelated.TLS and certificates for the HTTP API — a short page that is entirely about certificates, and never mentions rotation or replication.You can create the table and add the documents using the commands below.So, what we have is: one table, three vector columns with the same source text — one column per strategy. A single INSERT fills all three, so the comparison conditions are identical:DROP TABLE IF EXISTS docs; CREATE TABLE docs ( title text, body text, v_truncate float_vector knn_type='hnsw' hnsw_similarity='cosine' model_name='Xenova/all-MiniLM-L6-v2' from='title,body', v_mean float_vector knn_type='hnsw' hnsw_similarity='cosine' model_name='Xenova/all-MiniLM-L6-v2' from='title,body' chunk_strategy='mean', v_sentence float_vector_array knn_type='hnsw' hnsw_similarity='cosine' model_name='Xenova/all-MiniLM-L6-v2' from='title,body' chunk_strategy='sentence' max_tokens='128' overlap_tokens='32' ); Insert the four documentsINSERT INTO docs (id, title, body) VALUES (1, 'Backup and restore runbook', 'Nightly backups run at 02:00 UTC from the standby node. The job snapshots every table directory, writes a manifest, and uploads the result to object storage. Retention is thirty daily copies, twelve monthly copies, and one yearly copy. A restore drill runs on the first Monday of each month against a scratch cluster. The drill counts as passed only when a full-text search over the restored data returns the same document count as production. Anything less is treated as a failed drill and investigated the same week. Before a restore, freeze the target cluster so that no writes land while files are being replaced. Copy the manifest first and verify its checksum. If the checksum does not match, stop: a partial restore is worse than no restore, because the cluster will start and silently serve half the corpus. After the files are in place, unfreeze and let replication catch up. Watch the queue depth. If it does not drain within ten minutes, the node is probably still reading from cold storage and needs a warm-up pass before it can serve traffic. Backup failures page the on-call engineer. The three most common causes are an expired object storage credential, a disk that filled up while the snapshot was being written, and a table left frozen by a previous failed run. All three are recoverable without data loss. Check the job log first, then the disk, then the freeze state of every table. Capacity planning for backups is boring but it matters. A daily copy of the search cluster is roughly the size of the data directory plus fifteen percent for the manifest and metadata. Multiply by the retention count, add the transfer cost, and you have the monthly bill. Most teams discover too late that the yearly copies dominate the storage line. Object storage lifecycle rules do most of the retention work. Daily copies move to infrequent access after seven days and expire after thirty. Monthly copies move to archive after sixty days. Yearly copies never expire automatically; deleting one is a manual action that requires a second approver. Credentials for the backup job live in the secret manager and are issued to a role, not to a person. The role can write new objects and list the bucket. It cannot delete, and it cannot read objects older than the current day. That last restriction is the cheapest defence against a compromised backup runner turning into a data exfiltration path. Documentation for each table lives next to its schema: what the table is for, who owns it, how large it is expected to get, and whether it can be rebuilt from an upstream source. A table that can be rebuilt does not need thirty daily copies. Roughly half of most clusters turns out to be derived data that nobody had marked as derived. Verification is not the same as the job exiting zero. The job can succeed while producing an unusable copy: an empty table, a truncated upload, a manifest that references a file that was never written. The verification step reads the manifest back, checks every referenced object exists and matches its recorded size, and compares row counts on three sampled tables against production. Rotating the replication TLS certificate is a separate procedure and the step people most often get wrong. The certificate that secures the replication port is not the same as the one the HTTP API uses, and replacing one does not replace the other. Generate the new key and signing request on the node that will be rotated first, sign them with the cluster certificate authority, and place the files next to the existing ones rather than on top of them. Then update the node configuration to point at the new paths and reload. Do one node at a time and confirm that the cluster reports every peer as synced before moving on. A half-rotated cluster where two nodes trust different authorities will keep accepting writes on both sides and diverge quietly. When every node has been rotated, remove the old key material and revoke the retired certificate at the authority.'), (2, 'Monitoring and alerting guide', 'Every node exports metrics over an HTTP endpoint that a scraper collects once per fifteen seconds. The dashboards are grouped into four rows: traffic, latency, saturation, and errors. Traffic is queries per second broken down by table. Latency is the ninety-fifth and ninety-ninth percentile of query time, measured server side. Alerting is deliberately thin. Paging alerts fire on sustained error rate above one percent for five minutes, on ninety-ninth percentile latency above two seconds for ten minutes, and on a node dropping out of the cluster. Everything else is a ticket, not a page. Teams that page on every anomaly stop reading pages within a month. Log retention is fourteen days hot and ninety days cold. The query log records the query text, the table, the match count, and the elapsed time. Turning it on costs a few percent of throughput and is almost always worth it, because most performance investigations start with a slow query nobody knew was being issued.'), (3, 'Getting started with the CLI', 'The command line client connects over the MySQL wire protocol, so any MySQL client works and you do not need to install anything special. Point it at port 9306 and you get an interactive shell. The shell understands the usual conveniences: history, tab completion of table names, and vertical output when a row is too wide for the terminal. Start by listing tables, then look at one with SHOW CREATE TABLE. The output is the exact statement that would recreate the table, including every option that was applied implicitly, which makes it the fastest way to find out what a table actually does rather than what someone documented two years ago. Bulk loading from the shell is possible but rarely what you want. For anything above a few thousand rows, use the HTTP bulk endpoint or one of the log shipper integrations, both of which batch and retry for you.'), (4, 'TLS and certificates for the HTTP API', 'The HTTP API can be served over TLS. You supply a certificate, a private key, and optionally a chain file, and the listener starts speaking HTTPS instead of HTTP. Clients