---
title: On-premise chunking and document ingestion ⬣ POMA AI
description: Run POMA’s document ingestion and RAG chunking self-hosted, in a dedicated instance or in your own VPC, for data residency and regulated workloads.
canonical: https://www.poma-ai.com/on-premise-rag-chunking
generated: Markdown variant of the page above, built from the prerendered HTML
---
# On-premise deployment

On-premise chunking and document ingestion.

PrimeCut is POMA AI's document ingestion and chunking engine. It reads how a document is actually built, its headings, tables, captions and scanned pages, and splits it along those lines instead of at arbitrary character counts. Structure-aware chunks return the passage that answers the question and little else: on [our OfficeQA benchmark](https://github.com/poma-ai/poma-officeqa), 100% recall at 23% of the tokens. [Read more about what makes us different](https://www.poma-ai.com/products/primecut-rag-ingestion-chunking).

PrimeCut is also built for teams who cannot put source documents through a public API: regulated industries, disconnected networks, and data-residency rules a shared service does not satisfy. It does not have to run on our infrastructure. It can run in a dedicated instance of its own, inside your own VPC, or on hardware you operate yourself. Tell us the constraint and we will tell you which of the three deployments meets it.

[

Talk to us about on-premise

](https://meetings-eu1.hubspot.com/alex-nuss?uuid=8ecc05d3-e734-4091-bb5c-cccde61b0737)[

Try it yourself

](#on-premise-try)

## Why ingestion

Chunking is the stage where your documents are most exposed.

Retrieval runs on vectors and hashed tokens. Ingestion runs on the document itself: the full text, the tables, the scans. If there is one stage of a RAG pipeline worth pulling inside your own perimeter, it is this one.

-   ### It reads the whole document
    
    Structure detection works on the text as written: headings, tables, captions, scans. That stage needs the document, not a vector of it.
    
-   ### It produces what persists
    
    Chunks are what gets stored, embedded and searched for the life of the project. Where they are produced decides where they live.
    
-   ### It is the boundary worth arguing about
    
    A reviewer wants to point at the ingestion step and ask exactly what leaves it. On a scoped deployment that answer is short, written down, and agreed with you up front.
    

## Deployment options

Three places POMA can run.

The chunking engine is the same in all three. What changes is who operates the machine it runs on, and how far your documents travel to reach it.

Managed

### POMA cloud, EU-resident

Our hosted service, with processing and storage in the EU and a per-tenant cryptographic privacy layer over stored text, search vectors and keyword tokens.

-   Frankfurt and the Netherlands
-   Every subprocessor named and located
-   Free tier: 1,000 PrimeCut pages on sign-up

[

Read the security page

](https://www.poma-ai.com/security)

Dedicated

### Dedicated instance or VPC

POMA deployed in an instance of its own, or inside your virtual private cloud, with full data isolation and multi-user access control.

-   Full tenant isolation
-   Signed custom DPA
-   Dedicated technical support

[

Talk to us

](https://meetings-eu1.hubspot.com/alex-nuss?uuid=8ecc05d3-e734-4091-bb5c-cccde61b0737)

On-premise

### Your data centre, your hardware

A self-hosted deployment on infrastructure you own and operate, running entirely inside your perimeter, fully air-gapped where you need it. Try it on your own machine with a single docker pull; a production license comes through our team.

-   Try it with one docker pull
-   Runs fully air-gapped
-   Production license through our team

[

Talk to us

](https://meetings-eu1.hubspot.com/alex-nuss?uuid=8ecc05d3-e734-4091-bb5c-cccde61b0737)

## Compliance

What you can put in front of a security review.

We would rather hand a reviewer the list than a slogan. Every subprocessor on the managed service is named and located, along with the few steps where text is processed in the clear. Start there, then tell us which of those your policy will not allow.

Our full write-up, covering encryption at rest, per-tenant key isolation, and an honest account of the few steps where text is processed in the clear, lives on the [security page](https://www.poma-ai.com/security).

-   **Data residency.** Managed processing and storage run in the EU: compute in the Netherlands, search index and storage in Frankfurt. A dedicated, VPC or self-hosted deployment runs where you choose to put it.
-   **No training, anywhere.** Every provider we use is bound by a data-processing agreement with a no-training commitment and limited or zero retention.
-   **Per-tenant key isolation.** One secret per project, generated once, never reused across tenants, and rotatable on request.
-   **Signed DPA.** Custom data-processing agreements are part of the enterprise arrangement, as is multi-user access control.

## Try it yourself

Run it on your own machine in one command.

No sign-up and no sales call to evaluate it. The image runs unregistered, so you can point it at your own documents today. When you want it in production, the license comes through our team.

Ready for production? [Talk to us about a license](https://meetings-eu1.hubspot.com/alex-nuss?uuid=8ecc05d3-e734-4091-bb5c-cccde61b0737).

poma@localhost 

$ `docker pull registry.poma-ai.com/primecut:latest`

$ `docker run -d --name poma -p 8080:8080 \`

`-v ./documents:/data \`

`registry.poma-ai.com/primecut:latest`

$ `docker exec poma primecut ingest /data/msci_world_index.pdf`

## Common questions

On-premise, self-hosted and air-gapped.

### Can POMA run on-premise?

Yes. You can try it on your own machine with a single docker pull, no sign-up needed. A production license comes through our team, where we agree how the deployment is configured for your environment and what the support model looks like. Two nearer options are available today without that conversation: a dedicated instance, or a deployment inside your own VPC.

### Does POMA work in an air-gapped environment?

Yes. A self-hosted deployment can run fully air-gapped, with no outbound connection required. The ingestion steps that run through named external providers on our managed service (OCR on scans, image understanding, our fine-tuned chunking model) run inside your perimeter instead. One practical consequence: with no outbound route, critical security updates and license renewals are coordinated manually with our team rather than reaching you automatically. Tell us your constraints and we will confirm the configuration for your environment.

### Is POMA self-hosted or SaaS?

The default is a managed service with processing and storage in the EU. Teams that cannot use a shared service can run a dedicated instance or a deployment inside their own VPC, both of which give full tenant isolation and a signed data-processing agreement. A fully self-hosted install on your own hardware can be tried with a docker pull and licensed through our team.

### Where is my data processed on the managed service?

In the EU: compute in the Netherlands and the search index and storage in Frankfurt. Every subprocessor is named and located on our security page, each is bound by a data-processing agreement with a no-training commitment, and the few providers based outside the EU operate under a safeguard the EU recognises, such as an adequacy decision or standard contractual clauses.

### Does a self-hosted deployment change how chunking behaves?

No. The structure-aware chunking behaviour and the output format belong to the engine, not to the address it runs at, so a pipeline built against the managed API does not have to be rebuilt when it moves.

### What does a self-hosted deployment cost?

Trying it costs nothing: a docker pull runs it on your own machine. A production license is agreed with our team rather than listed per page, because the cost depends on scale and the support model. The managed service and its per-page pricing are on the pricing page; dedicated instances and VPC deployments sit under the enterprise tier there.

## Next step

Start where it is easiest.

The cheapest way to find out whether the retrieval is good enough is to run it on real documents first, on the free tier, and take the deployment question to your security team with an answer in hand. The chunk format does not change with the address, so nothing downstream has to be rebuilt if you move later.

Read up on the engine itself on the [PrimeCut page](https://www.poma-ai.com/products/primecut-rag-ingestion-chunking), our managed RAG SaaS, [index4.ai (opens in a new tab)](https://index4.ai), or the tiers on [pricing](https://www.poma-ai.com/pricing).

[

Talk to us about on-premise

](https://meetings-eu1.hubspot.com/alex-nuss?uuid=8ecc05d3-e734-4091-bb5c-cccde61b0737)[

Try the managed service free

](https://console.poma-ai.com/?mode=register)

1,000 free pages on sign-up. No credit card required.
