Source: http://www.poma-ai.com/docs/learn/ingestion/ingestion-patterns

# Ingestion patterns

Once you know ingestion is more than just "read file, get text," the next question is how to wire it into the rest of your stack. In practice, teams converge on a few common patterns.

## File-based batch ingestion

This is the default approach for many early RAG deployments: periodically ingest documents from shared folders, buckets, or repositories as batches.

**Typical use cases**

- Periodic ingestion of internal knowledge bases and policy documents.
- Migrating legacy archives of PDFs and Word files into a new RAG system.
- One-off ingest of open-source document sets such as manuals and standards.

**Advantages**

- Simple operational model.
- Works well for large document types where real-time updates are not required.
- The normalized corpus can be reused with different embedding models later.

**Disadvantages**

- Poor fit for real-time use cases.
- Coarse-grained error handling.
- Updates and deletions are harder to reconcile cleanly.

## API- and event-based ingestion

Here, document ingestion reacts to events. A new ticket is created, a wiki page is updated, or a file is uploaded through an application. The ingestion pipeline is triggered via API, queue, or webhook.

**Typical use cases**

- Customer support systems where new tickets should become searchable within seconds.
- Product docs that must reach RAG-powered chat quickly after an update.
- Workflow tools that embed RAG inside an existing SaaS product.

**Advantages**

- Supports near-real-time updates and deletions.
- Gives you finer-grained routing and metadata control.
- Lets you vary parsing strategies by source or format.

**Disadvantages**

- Operationally more complex.
- Harder to rebuild from scratch after systemic changes.
- Easier to couple tightly to upstream producers.

## Connector-based ingestion

Many RAG stacks rely on connectors to extract content from SaaS platforms or transactional databases and map it into a neutral internal representation.

**Typical use cases**

- Building organization-wide search across many systems.
- Pulling tickets, CRM data, and knowledge-base content into one retrieval layer.
- Standardizing authentication, pagination, and rate limits through shared integrations.

**Advantages**

- Reduces implementation time.
- Often aligns initial structure with business semantics such as tickets, issues, or wiki pages.
- Makes multi-system ingestion easier to centralize.

**Disadvantages**

- Limited control over parsing fidelity.
- Potential vendor lock-in.
- Not every connector exposes enough structural detail to drive good chunking.

## What each pattern has to hand to the chunker

The pattern decides how documents arrive. It also decides what the chunker gets to see, and that is where most retrieval problems are actually created.

| Pattern | What tends to survive | What tends to get lost | What to preserve deliberately |
| --- | --- | --- | --- |
| File-based batch | The original file, so a structure-aware parser can recover headings, tables and page numbers. | Provenance: which folder, which version, when it changed. | A stable document ID and version per file, so re-ingests replace rather than duplicate. |
| API- and event-based | Freshness and per-source metadata such as author, ticket ID, timestamps. | Structure, because payloads are often already flattened to plain text by the producer. | Send the source format (HTML, Markdown, the original attachment) rather than a text field, or the chunker cannot rebuild the hierarchy. |
| Connector-based | Business semantics: page, issue, record, and their relationships. | Formatting fidelity; many connectors return a normalized text blob. | Ask the connector for rendered HTML or the export format, and keep the parent-child relations as metadata. |

Two rules hold across all three:

- **Chunk from the richest representation you have.** A PDF parsed with layout detection yields headings and tables; the same content pasted as `text` yields neither. If a producer can only send flat text, expect flat chunks.
- **Keep the document identity through the whole path.** Updates and deletions are only possible if every chunk points back to the document and version it came from. Batch pipelines lose this by re-ingesting under new IDs; event pipelines lose it when the event carries the content but not the source.

## Combining patterns

In mature deployments, teams combine all three: batch for static archives, event-based flows for live content, and connectors for the long tail of SaaS systems. The pattern mix is an operational choice. The chunking step should be the same for all of them, so that a policy PDF from the archive and the same policy pasted into a wiki page produce comparable chunks. That is why the chunking step is worth centralising. PrimeCut does that from the original file when it has one, or from a parser's output when it does not.

<Tldr>
Batch, event-based, and connector-driven ingestion all have legitimate use cases. The important design requirement is that the pattern you choose still preserves enough structure for chunking and retrieval later.
</Tldr>

## Continue reading

- [Tooling comparison](/learn/ingestion/tooling-comparison) — how different tools handle ingestion
- [System design](/learn/ingestion/system-design) — designing ingestion and chunking as one system
- [The full ingestion guide](/document-ingestion-chunking-rag) — complete narrative guide