Source: http://www.poma-ai.com/docs/pipelines/paddleocr-vl-to-qdrant

# The Missing Link Between PaddleOCR-VL and Optimal Retrieval in Qdrant

<ByAuthor />

**The short answer:** PaddleOCR-VL gives you self-hosted, layout-aware OCR — PP-DocLayoutV3 labels every region, PaddleOCR-VL reads it, all behind your own HTTP endpoint. Qdrant gives you excellent hybrid dense+sparse vector search. Wired together naively — flatten, split, embed — they still produce mediocre RAG, because nothing in between rebuilds the document's hierarchy or respects PaddleOCR-VL's own reading-order field. The missing link is POMA: `PrimeCut().ingest()` consumes the raw layout-parsing JSON (auto-detected, either accepted shape), emits hierarchy-preserving [chunksets](/learn/chunking/chunksets), and `PomaQdrant.upsert_poma_points()` writes them as hybrid dense+sparse points with indexed hierarchy payloads.

## What each end of the pipeline actually provides

**PaddleOCR-VL** runs as a self-hosted pipeline — PP-DocLayoutV3 for layout detection, PaddleOCR-VL served via vLLM for reading, behind a `/layout-parsing` HTTP API you operate yourself. It returns either the PaddleX serving envelope (`result.layoutParsingResults[]`, each with a `markdown` object and a `prunedResult` block list) or a raw `save_to_json()` page dict with `parsing_res_list`: flat blocks carrying `block_label`, `block_content`, `block_id`, and `block_order`. What it doesn't provide: an opinion on what a retrieval unit should be, or a link between a `paragraph_title` on page 12 and the `doc_title` on page 1. Details: [The Optimal Chunker for PaddleOCR-VL](/optimal-chunker-paddleocr-vl).

**Qdrant** provides named dense + sparse vectors on one point, payload indexes for filtered ANN, and HNSW with quantization options — open source or Qdrant Cloud. What it doesn't provide: any opinion about what a point should contain. It retrieves nearest neighbors of whatever you embedded, including scrambled or context-free fragments, if that's what you gave it. Details: [The Optimal Chunks for the Best Retrieval in Qdrant](/optimal-chunks-qdrant).

## The naive wiring, and where it breaks

The common recipe — call `/layout-parsing`, concatenate `block_content` (sorted by `block_id`, the field that happens to come first in the JSON) or join `markdown.text` across pages, run a `RecursiveCharacterTextSplitter`, embed, `client.upsert(...)` — breaks in a pair-specific way:

- **Sorting by `block_id` instead of `block_order` scrambles reading order.** `block_id` is assignment order, not reading order; on a multi-column page the two diverge. A pipeline that sorts the wrong field embeds a logically reordered document, and Qdrant's BM25 sparse vectors then represent clause sequences that never appeared in that order in the source.
- **PaddleOCR-VL's `page_index` vanishes at the join**, so no Qdrant payload can cite a page, and per-document filters are all that's left.
- **Overlap inflates the collection.** Splitter overlap embeds every boundary span twice, so top-k results arrive as near-duplicates of each other, crowding out the passage that actually answers the question — regardless of how clean the underlying OCR was.

Chunksets fix this at the unit level: `block_order`-respecting assembly, no overlap, page and hierarchy metadata intact. On a reference legal document, this approach answered the same query with 337 tokens of retrieved context versus 1,542 for a recursive character splitter, with zero information loss ([methodology](/document-ingestion-chunking-rag)).

## The pipeline, end to end

```bash
pip install requests 'poma[qdrant]'
```

```python
import json
import os
import requests
from poma import PrimeCut
from poma.integrations.qdrant import PomaQdrant

# 1. Your existing self-hosted call — unchanged. `/layout-parsing` is the
#    PaddleX-compatible endpoint your layout-server + PaddleOCR-VL stack exposes.
resp = requests.post(
    "http://layout-server:40111/layout-parsing",
    json={"file": "<base64-encoded PDF>", "fileType": 0},
)
resp.raise_for_status()
with open("contract.paddleocr-vl.json", "w") as f:
    json.dump(resp.json(), f)

# 2. The missing link — raw result JSON in, chunks + chunksets out.
poma = PrimeCut()  # reads POMA_API_KEY
result = poma.ingest("contract.paddleocr-vl.json")  # PaddleOCR-VL shape auto-detected

# 3. Qdrant — hybrid points with hierarchy payloads, one call.
qdrant = PomaQdrant(
    url=os.environ["QDRANT_URL"],
    api_key=os.environ["QDRANT_API_KEY"],
    cloud_inference=True,
    collection_name="contracts",
    dense_model="sentence-transformers/all-MiniLM-L6-v2",
    sparse_model="Qdrant/bm25",
    dense_size=384,
    auto_create_collection=True,
)
qdrant.upsert_poma_points(result)

# 4. Retrieve prompt-ready context.
cheatsheets = qdrant.get_cheatsheets(
    query="What are the early termination conditions?",
    limit=3,
)
print(cheatsheets[0]["content"])
```

POMA validates the payload shape up front (a corrupted or mislabeled upload 422s immediately), prefers `markdown.text` when present and falls back to `parsing_res_list` sorted by `block_order` when it isn't, drops running furniture (`header`/`footer`/`page_number`/`aside_text_number`), and rebuilds the cross-page heading tree before chunking. To force or suppress detection, set `external_ocr_source` to `"paddleocr_vl"` or `"none"`.

## Metadata mapping: POMA fields → Qdrant primitives

| POMA chunk field | Qdrant primitive | What it enables |
| --- | --- | --- |
| `to_embed` | dense vector + BM25 sparse vector (named vectors, one point) | paraphrase + exact-term hybrid retrieval |
| `file_id` | payload field, **payload-indexed** | scope queries to one document |
| `page` (from PaddleOCR-VL `page_index`, 1-based) | payload field, **payload-indexed** | page-cited answers, page-range filters |
| `depth` | payload field | filter/re-rank by hierarchy level |
| `chunk_index` | payload field | stable ordering, independent of `block_id` |
| chunkset lineage | payload (`chunk_details`) | cheatsheet assembly without a second store |

## Frequently asked questions

### How do I get self-hosted PaddleOCR-VL results into Qdrant for RAG?

Save the raw `/layout-parsing` JSON, run `PrimeCut().ingest()` on it (auto-detected, hierarchy rebuilt), then `PomaQdrant.upsert_poma_points(result)` — hybrid points with `file_id`/`page`/`depth` payloads. Retrieve with `get_cheatsheets(query=...)`.

### Why does block_order matter more than block_id when feeding Qdrant's hybrid search?

`block_id` is detection order, not reading order; `block_order` is. Sorting by the wrong field scrambles multi-column text before it's embedded, so Qdrant's BM25 sparse vectors represent clause sequences that never occurred in that order in the source.

### What Qdrant payload fields should PaddleOCR-VL chunks carry?

`file_id`, `page`, `depth`, `chunk_index` plus content; payload-index `file_id` and `page`. `PomaQdrant` writes these by default regardless of which accepted PaddleOCR-VL shape produced them.

### Can I keep documents on my own infrastructure through both OCR and retrieval?

Yes — PaddleOCR-VL and Qdrant can both run self-hosted. Only the layout-parsing result JSON needs to reach POMA's API for chunking; the source document and the vector index never have to leave your infrastructure.

### Do images in the PaddleOCR-VL result survive the trip to Qdrant?

Yes, when present: inline base64 or referenced bytes are described and embedded as searchable text. Empty image blocks become a visible `[IMG-N]` marker and are counted — never silently dropped.

## Related recipes

Same parser, different store: [PaddleOCR-VL → Pinecone](/pipelines/paddleocr-vl-to-pinecone) · [PaddleOCR-VL → Weaviate](/pipelines/paddleocr-vl-to-weaviate) · [PaddleOCR-VL → Chroma](/pipelines/paddleocr-vl-to-chroma)

Same store, different parser: [Mistral OCR → Qdrant](/pipelines/mistral-ocr-to-qdrant) · [LlamaParse → Qdrant](/pipelines/llamaparse-to-qdrant) · [Docling → Qdrant](/pipelines/docling-to-qdrant)

Foundations: [The Optimal Chunker for PaddleOCR-VL](/optimal-chunker-paddleocr-vl) · [The Optimal Chunks for Qdrant](/optimal-chunks-qdrant) · [All pipeline recipes](/pipelines/) · [RAG architecture guide](/guides/rag-architecture/)