Source: http://www.poma-ai.com/docs/learn/chunking/strategy-landscape

# Which chunking strategy for which document?

The [full chunking guide](/rag-chunking-strategies-text-splitters) describes every strategy in depth. This page answers a narrower question: given the documents you actually have, which family should you start from, and what will still go wrong.

## Start from the document, not the tool

Chunking strategies differ in one thing: how they choose boundaries. Whether a boundary rule works depends far more on the shape of your documents than on the rule itself. Three questions sort most corpora:

1. **Does the text carry usable structure?** Headings, numbered sections, tables, code blocks. Contracts, manuals, standards and technical documentation usually do. Chat logs, transcripts and scraped prose usually do not.
2. **How long is a typical document?** Under a page, whole-document embedding is often enough. Over ten pages, boundary choice starts to dominate retrieval quality.
3. **What do questions look like?** Factoid lookups reward small, precise chunks. "Explain the termination conditions" rewards chunks that keep their surrounding context.

## Decision table

| Your documents | Start with | Why | What still breaks |
| --- | --- | --- | --- |
| Short notes, tickets, FAQs (under a page) | [No chunking](/rag-chunking-strategies-text-splitters#no-chunking) | One embedding per document already fits the model window and preserves everything. | Long tail of oversized items; retrieval returns the whole item even for one-line answers. |
| Plain prose without headings (articles, transcripts) | [Recursive delimiter](/rag-chunking-strategies-text-splitters#recursive-delimiter-chunking) or [sentence chunking](/rag-chunking-strategies-text-splitters#sentence-paragraph-chunking) | Paragraph and sentence boundaries are the only structure available, and these rules respect them. | Topic drift inside a chunk; answers that depend on a sentence two paragraphs back. |
| Prose where topics shift without visible markers (research notes, meeting minutes) | [Semantic similarity chunking](/rag-chunking-strategies-text-splitters#semantic-similarity-chunking) | Boundaries follow meaning rather than punctuation. | Extra embedding cost at index time; boundaries are still cuts, so context outside the chunk is lost. |
| Markdown, HTML, code | [Format-structure chunking](/rag-chunking-strategies-text-splitters#format-structure-chunking-markdown-headings-html-tags-code-blocks) | The format already marks sections and blocks; splitting on them keeps units coherent. | Sections that exceed the chunk budget fall back to size-based cuts; heading context is not attached to the chunk. |
| PDFs with tables, figures and nested sections (contracts, manuals, filings) | [Partition-then-pack](/rag-chunking-strategies-text-splitters#partition-then-pack-element-based-chunking) or [hierarchical chunking](/rag-chunking-strategies-text-splitters#hierarchical-chunking) | Elements are detected first, so tables and captions stay whole; hierarchy lets you retrieve coarse or fine. | Parser quality caps everything downstream; parent-child schemes add index complexity and still return isolated units. |
| Very long documents with cross-references (standards, legal codes) | [Hierarchical chunking](/rag-chunking-strategies-text-splitters#hierarchical-chunking) plus retrieval that returns ancestors | Questions span levels; a clause is meaningless without its section. | Most implementations retrieve a leaf and hope the parent is enough; the full path is rarely reconstructed. |
| Anything where recall must not miss boundary facts | [Late chunking](/rag-chunking-strategies-text-splitters#late-chunking) | Token embeddings see the whole document before pooling, so a chunk's vector reflects its neighbours. | Requires long-context embedding models; the returned text is still a slice. |

## The failure that survives every row

Every strategy in the table still returns a slice of a flat sequence. The better ones choose better cut points, but the retrieved unit arrives without the trail that says what it belongs to. That is why chunk size and overlap tuning plateaus: the parameters move the boundary, they do not remove it.

[POMA chunksets](/learn/chunking/chunksets) change the unit instead of the boundary. A chunkset is a complete root-to-leaf path through the document, so every retrieved sentence carries its chapter, section and paragraph context. That is the approach behind PrimeCut, which parses the document itself or takes the output of a parser you already run.

<Tldr>
Match the strategy to the document: no chunking for short items, recursive or sentence rules for plain prose, semantic when topics shift invisibly, structure-aware and hierarchical for anything with headings and tables. Then plan for the failure they share, which is context lost at the boundary.
</Tldr>

## Continue reading

- [Common failure modes](/learn/chunking/common-failure-modes) — why even advanced strategies lose context
- [POMA chunksets](/learn/chunking/chunksets) — the non-breaking alternative
- [Strategy comparison table](/learn/chunking/strategy-comparison) — all strategies side by side
- [The full chunking guide](/rag-chunking-strategies-text-splitters) — every strategy in depth, with chunk size and overlap guidance